
Share
Podcast
About This Episode
Enterprises have spent decades writing governance policies that quietly assumed a human would exercise judgment at the final step. Jake Williams, VP of R&D at Hunter Strategy and a faculty analyst at IANS Research, argues that assumption collapses the moment autonomous agents enter the picture. Agents do not use discretion and cannot be held accountable the way an employee can, which turns the implied trust at the bottom of every access control list into a liability. The result is a widening gap between what organizations govern on paper and what they can actually enforce at machine speed.
Williams makes the case that the legacy problems security teams tolerated for years, coarse data classification and unresolved entitlements debt chief among them, are now being exposed by the speed and scope of AI-driven activity. He walks through why volumetric detection matters more than signatures against agent-scale behavior, why every deployed agent needs a named business owner who is accountable for its actions, and how his open framework CUSTODY gives teams a shared taxonomy to classify and contain agents before they ship. The conversation closes on the economics most organizations are avoiding, from the true cost of tokens to the 15-30% security premium of screening for prompt injection.
For readers who want to go deeper, Williams references 1Password research on why AI-generated vulnerability patches still require expert human review and a Microsoft Entra analysis of why OAuth must evolve to support AI agents.
Podcast
Popular Episodes

35 mins
The War on Data, Cyberspies and AI with Eric O'Neill - Part I
Episode 352
January 6, 2026

22 mins
How AI and Third-Party Risk Are Transforming Healthcare Cybersecurity with Ed Gaudet - Part I
Episode 339
October 1, 2025

24 mins
The Evolving Cyber Threat Landscape in Healthcare: Insights from Fortified Health Security’s Russell Teague - Part I
Episode 332
August 11, 2025

44 mins
From Battlefield to Boardroom: Ricoh Danielson’s Lessons on Cyber Warfare and Digital Forensics
Episode 327
May 27, 2025
Podcast

[00:00] Welcome, Jake Williams
Rachael Lyon:
Welcome to the To The Point Cybersecurity Podcast. Each week, join Jonathan Knepher and Rachael Lyon to explore the latest in global cybersecurity news, trending topics, and cyber industry initiatives impacting businesses, governments, and our way of life. Now, let's get to the Point.
Hello, everyone. Welcome to this week's episode of To The Point podcast. I'm Rachael Lyon, here with my co-host, Jon Knepher. Hi, Jon.
Jonathan Knepher:
Hello, Rachael.
Rachael Lyon:
What's been going on? I feel like we haven't caught up in a century.
Jonathan Knepher:
I know it's been a while, but I am so glad that the weather is cooling down. It was so brutal here, like nothing could keep up with the heat.
Rachael Lyon:
In San Diego. Okay.
Jonathan Knepher:
In San Diego. It was 105.8 at my house for a couple of days, which was too much for me.
Rachael Lyon:
But it was a dry heat or no?
Jonathan Knepher:
Oh no, it was humid this time.
Rachael Lyon:
It was a humid one, okay.
Jonathan Knepher:
With the whole El Niño precursor stuff coming in.
Rachael Lyon:
Oh yeah, that is too bad. That is too bad. I'm ready for fall. We don't really get fall here in Texas. I think probably similar to you, but I'm here for it. I'm ready. I'm ready, ready. Without further ado, let's jump into this week's guest.
Excited to welcome Jake Williams. He's known in the security community as Malware Jake. He's VP of R&D at Hunter Strategy and a faculty analyst at IONS Research, where he advises CISOs at organizations ranging from the Fortune 10 to mid-market. He's worked both sides of the keyboard. A former NSA-designated Master Network Exploitation Operator, he spent years in offensive operations and incident response before turning to advisory work. And he's a 2-time winner of the DOD Cybercrime Center's Digital Forensics Challenge. His current research sits where AI and machine learning collide with security, including his open framework for containing autonomous agents, which we'll talk about a little later. And he also co-hosts the Breach Please podcast.
Welcome, Jake.
Jake Williams:
Hey, so glad to be here. It's fantastic.
[02:09] AI Turns Old Tech Debt Into New Exposure
Jonathan Knepher:
Yeah, thanks, Jake. Well, let's kick it off with the exciting bit, right? If you were still in offensive, what tools, what processes would you be using today?
Jake Williams:
Woof. No question I'd be using AI. I mean, I think back to, you know, even using like Ansible playbooks back in the day, right? The whole idea of like automating attacks is not new. What is new is really the ability to use, you know, something like generative AI to interpret data. I mean, certainly again, we were doing a lot of the pattern matching and whatnot, right? But unstructured data, right, that's something I couldn't do, right? And now having that ability at my fingertips, yikes, I, ah man, we could scale so much faster.
Jonathan Knepher:
So given that scale though, what can all of us on the defensive side do to protect against that?
Rachael Lyon:
Uh, woof.
Jake Williams:
You know, I think really here the unfortunate answer is going to be and it is an unfortunate answer, I think is going to be that we have to really look at security tech debt, right? So we've had a lot of technical debt over the years that hasn't been a big deal, entitlements management being a big one. And, you know, now it is a big deal, right? So because of the speed of AI attacks, you know, ultimately that tech debt is coming home to roost. So this is one of the spots where I tell folks regularly, like, you know, I'm not worried about the brand new agentic attacks per se, or the AI attack surface so much as I am the legacy attack surface that's now exacerbated by the speed of AI agents.
Rachael Lyon:
Yeah.
Jake Williams:
Not just speed, scope, right? Because, like, you wouldn't have necessarily put, you know, checking directory permissions on every subdirectory in the, you know, in the network, trying to find a place you can, you know, add a DLL to get DLL sideloading with a privileged account. Like, that's tedious work. It's just tedious work, right? It turns out it doesn't matter. The agents will do it, right? The agents don't care about tedium.
Jonathan Knepher:
And they'll find it.
Jake Williams:
Yeah, absolutely. Yeah, yeah. What I am telling folks is volumetric detections for the win, right? You may not know what, right? Traditionally, we've done signature-based primarily. A lot of my more advanced clients already had some volumetric detections where they looked for sometimes 2 or 3 standard deviations out of the mean, right, for a particular log entry. I mean, either direction, right? In a low direction, you mentioned an SRE, right? Oftentimes security can detect something's broken, right? Because we're not getting the log volume we expect. But now on the agent side, though, it's that explosive log volume. Think of the Hugging Face attacks, right?
Rachael Lyon:
Right.
Jake Williams:
From OpenAI. If they'd had volumetric detection in place, they'd have seen it immediately. Same thing with RubyGems.
[04:59] Why Data Classification Can't Govern Agents
Rachael Lyon:
Sorry, I'm writing all these little nuggets down. So, Jake, what's really interesting to me right now is governance is kind of becoming a sexy word, if I could say that. And it's kind of a really fascinating topic, particularly in this realm of, you know, agentic enterprises, etc. What's at the core of why the gap— there's a gap between governance and enforcement, like, how would you break that down?
Jake Williams:
Ooh, that's really good. That's a good question. You know, I'm doing a lot of AI work with some large financial services companies, a couple of other sectors as well. And one of the first things that we do is, you know, we do talk governance and people are like, no, no, no, I want to talk about technical controls. And I'm like, brother, I can't apply a technical control until you tell me what you want to do in the first place. One piece of governance that is not here to save us, like counter to the traditional wisdom as it were, is data classification. I am yet to see a data classification policy, and this isn't just because it's not applied universally. We have a decade of experience plus showing us that this is a, I'll say, a losing battle for sure, but let's call it close, right? And so when we look at most folks' data classification policies and they think about applying labels with those classifications to different files, I say, okay, well, you have your sensitive internal use only.
Can a particular AI system use this, yes or no? And they're like, well, some of the data. I'm like, well, then this is useless as a control, right? Because if the answer is sometimes, right, I can't codify sometimes into a, you know, effectively into a governance rule. And so we're finding there is even something as basic and as mundane as data classification, we're having to really back up, right? And really back up and say, is this granular enough to support— again, hate to say the G word again, right, back to the governance— is it granular enough to support the governance requirements, right, around this data? And that's just the start of it, right? We could spend all day talking identity governance and, you know, how that's applying to agentic work, but just even get the data governance right is a challenge for the vast majority of organizations. And again, that's before we layer AI on top of anything. This is just asking the question of, how will we implement these technical controls? What are your desires around these?
Jonathan Knepher:
But I feel like you still need that data classification, right? Like, you can't win the battle without it, right? Right.
Jake Williams:
I'm not saying that we shouldn't do data classification. What I'm saying is that for the vast majority of organizations, their classification labeling schema is too coarse, right? It's not fine-grained enough, right, to be able to be used for policy decisions. And if I can't use it for a policy decision, then, you know, ultimately I can't use it for effective data governance for agents, right? The ideal would be I could look at a particular piece of data, I could look at the policy label, and I could say, yes, this can be processed by this AI, or no, this can't. And with most organizations having, you know, a 3, 4, or in some cases 5-tier, you know, data classification standard, you know, as you get into the sensitive internal use, you know, typically that 2nd or 3rd tier, as we get into that tier, Mm-hmm. The answer is sometimes, right? And again, I just can't codify sometimes into a technical control, right?
Rachael Lyon:
So how does that, I mean, ostensibly apply to when you're adding policy, right? I mean, how do you account for these exceptions? Because, right, I mean, access to sensitive data is critical for AI agents to actually bring value to an organization. But to your point, you know, is it a litany of if this, then that, and/or, I mean, you know, kind of what is the way forward for trying to manage through this?
Jake Williams:
I mean, it can be ultimately, you know, some set of conditionals, you know, Boolean conditionals, more advanced logic. Obviously, the more advanced the logic is, the harder it is to maintain. Forget the setup cost, that's a one-time CapEx, or I look at as a one-time CapEx cost, but then I'm also looking at what does it cost to maintain it? What's the OpEx look like long-term? And of course, Jon being an SRE man, and please, I'm not picking on SREs, right?
Jonathan Knepher:
Oh, pick on it. It's all good.
Jake Williams:
Look, from an incident response standpoint, right, I think you know that a non-insignificant number of incidents that I work, right, the root cause was an SRE making something work again, right? That later then becomes the incident, right? So the more complexity is the enemy of security. It always has been, it always will be. So, you know, the solution, you know, Rachael, coming back around, I think the solution is to get more fine-grained with your data classification schema.
Jonathan Knepher:
Mm-hmm.
Jake Williams:
Now, obviously, again, that's adding complexity to a different layer, right? Do we apply the right classification label here? But the reality is that most data classification schemes, again, were designed with humans in mind. And what I always tell folks is that those classification schemas and thinking about DLP, that's there to help the human. And as we think about entitlements, there's an implied rule at the bottom of any entitlements list that says Jake or Jon or Rachael will use their judgment. They won't do something stupid with this data. And if they do, we're going to fire them, we're going to counsel them, we're going to criminally charge them for criminal activity, criminal prosecution, whatever. And you can't do that with an agent, right? So that implied, you know, bottom of the access control list that we've been dealing with now, living with for decades, suddenly, like, that now becomes a liability, much more so than it was before.
[10:40] The Named Owner Who's Accountable for the Agent
Jonathan Knepher:
So given that, where in the organization should kind of ownership of this enforcement and policy Does this need to be independent in like the CISO role? Should it be an SRE? Like, where should it be?
Rachael Lyon:
Woof.
Jake Williams:
My Konami cheat code, if you will, right, for, you know, thinking about agentic permissions and entitlements, this does start at governance. And it's that we implement a policy that for each agent, right, that we're deploying, every agent that we're deploying, there's a named business owner who is accountable for the actions the agent takes. And where I've gotten that deployed, I don't have it deployed everywhere or enforced everywhere, but in the organizations where I've gotten that enforced, we start having adult discussions. Forgive the term here, but adult discussions about security versus, you know, otherwise what I find is I've got the business unit over here saying enable this agent and it's going to save us X dollars. And we can talk about does that really save us dollars or not? That's a whole other discussion, right? we could spend days on, but they are talking in terms of productivity or functional enablement, and I am talking security. Unfortunately, unlike the notebooks we used to have in grade school, like the converter between metric and the English system, I do not have a conversion for how to convert security into dollars, security into enablement. So what I try to do here then is make sure that the business unit is also talking security. And when, you know, we turn around and say, hey, you will be responsible and ultimately accountable for the actions that this agent takes, let's talk about security controls.
Let's think about least privilege, right? What is it you're actually asking for? And maybe one other cheat code that I'll throw out here, you know, for like, where does this live? Because that puts it back to the business unit. I'm gonna do ultimately the controls for you, right, as you lay them out and say, we accept this risk. risk. What I can't do in a security organization, I can't take on the liability of, you know, the business saying go, and then you and I both know cybersecurity is gonna be the poop umbrella for the org. If it touches a computer and nobody else says, I've got it, right? Yeah. It's fallen to us, right? So, yeah.
Jonathan Knepher:
So, on this ownership bit though, right? Like, I found it's really hard to, like, keep up with velocity, right? Like, Like, your engineers and your operators are under so much pressure now for efficiency and productivity. And the agents that we use want to do what you're asking them to do, kind of at whatever cost, it seems like. And, you know, reviewing literally every action and every line of code that they output doesn't seem feasible for a lot of organizations. Like, how do you reconcile that?
Jake Williams:
Well, again, this is, I think, where we have to come back and ask, like, you know, again, what risks are we ready to accept, right? And then align those risks with the utility that we're getting from the AI. And of course, that utility has to be balanced against cost, right? It becomes another kind of pillar in that. But your point about reviewing every line of, you know, every line of code, I believe it was 1Password that did the study recently. Did you see that with the— where they looked at AI agents and how well they patched vulnerabilities in code?
Jonathan Knepher:
I didn't see that, actually.
Jake Williams:
Oh, this is amazing. Let me see if I can pull the link so I'm not— yeah, here we go. Off-By-1 Labs. So ultimately here, they went through and did a systematic review of— and I'm going to drop it in chat here so you've got it as well. Let's see. There we go. Message all participants. So they did this fantastic study here where they actually went and looked at specific vulnerabilities and looked at, you know, how often does the agent correctly patch a vulnerability.
Now, I don't mean to imply that this is, you know, like representative of all agentic coding, right? But it is a place where we can objectively say the quality of the— where we can objectively measure the quality of the output. And when they did that, they found that it only patched the vulnerability correctly without some other side effect 26% of the time.
Jonathan Knepher:
Oh, wow.
Jake Williams:
This is like, listen, as an incident responder who bills by the hour, rock on, keep doing this, right? I'm loving this, right? The second biggest category, you know, ultimately was 49%, right? Beyond the 26% that it patched correctly, 49%, it patched the vulnerability, but it introduced behavioral changes in the application. Now, 1Password didn't go through, and they're Off-By-1 Labs, didn't go through and categorize in that, because that's the majority case, right? 49%. They didn't go through and categorize how many of those behavior changes introduce a vulnerability, but I'm going to tell you, a huge number of them do, right? I mean, we see that all the time. State changes within web applications or even just regular apps, what have you. And then they had others where they were able to see conclusively using their SAST system, SAST scanning afterwards, that Yes, we patched the vulnerability, but then we also introduced a brand new one. It got picked up in SAST.
Jonathan Knepher:
Yeah.
Jake Williams:
And so, like, you know, if I take that and then I extrapolate that back to, okay, this is just vulnerability fixing. Again, this is something that we have an oracle to measure against, right? To say, you know, that the output is functionally correct or not. Man, how do you review every line of code? I don't know that you can, right? I think you then have to accept that risk. Several of my clients, by the way, quick side note here, several of my clients who have tracked, you know, since the early days with GenAI coding, they're able to see, yes, they're shipping features faster, but they've also noticed a higher defect rate and a much higher time to resolve each defect. All of this aligns well with my developer experience, right? Because, you know, if I didn't write the code, then I'm looking at a codebase that I didn't write. It's just like putting a brand new developer right in a codebase and saying, go fix it.
Rachael Lyon:
Wow, there's so much to unpack there.
Jake Williams:
Sorry.
[16:50] Navigating Agentic AI When Best Practices Don't Exist Yet
Rachael Lyon:
I love it. I'd love to dig in a little bit on your current research on agentic AI and security. Obviously, agentic AI creates a vastly new kind of threat ecosystem, threat model. I mean, how do organizations even start to navigate this when we're working with the unknown, right? There are no best practices for the unknown today, and anything you do today is likely going to change next week. I mean, how do you move forward? Yeah.
Jake Williams:
Well, so the first thing that I'm working with, you know, organizations on is stakeholder expectation management, right? I think that's every bit as important as security because it lays the foundation for security enablement. And so, I typically am, you know, ensuring that as we're briefing up to stakeholders, we are explaining a couple of things. First off, if you want to run fast, great, right? You want to take advantage. You don't want to be left behind. Understand, though, we are building best practices literally, right, as you are deploying this stuff. And those best practices, to the extent that every security best practice is something that is fluid, these are particularly fluid. And so we have to then rethink how we basically, you know, estimate CapEx, you know, for deploying those best practices. Because I, you know, in something moving that fast, where we know best practices are shifting under our feet, I don't think it's appropriate to say, hey, one-time CapEx and then OpEx will take care of it, right? You know, because typically CapEx, we're bringing in consultants, right? We are bringing in, you know, surge staff.
And then, okay, now it's steady state. Okay, we can manage that. But that's not the reality here.
Rachael Lyon:
Right.
Jake Williams:
The second thing that I'm highlighting is that to the extent that any security tool, observation platform, etc., is a bet, right? Any deployment is a bet. Now, usually they're pretty good bets. We've got some level of history about the company. We've got some level of history about the technology that they're securing, and we can look across, you know, peers, industry analyst reports, what have you, and really get a pretty good feel for, yeah, this is a good bet. This is likely to, you know, be the correct EDR, SIEM, API gateway, etc., for us. This field is moving entirely too fast for that. So I'm briefing back to my CFOs, you know, my stakeholders who are approving budget, that I'm making some gambles here to support you moving fast, right? And I'm going to get a lot more of those wrong, right? I'm trying to make sure that I've seeded that, you know, seeded that understanding, right, ahead of there being some really big misunderstandings down the road. Like, Jake, you're wrong about that.
And I'm like, I'm flying blind trying to support you, brother. Right?
Rachael Lyon:
Right.
Jake Williams:
So, yeah.
Jonathan Knepher:
So I guess what all should we be watching and focusing on? Like, you've left a lot of open questions.
Jake Williams:
Yeah. Well, I mean, so what should you be focusing on? Yikes. You know, every— I rarely speak in absolutes, but I can speak here and say that every AI deployment that has been slowed down by security, and slowed down meaning like the business looked like, whoa, gotta pause this, right? Every single one of them has boiled down to one of two things. It's either data security or it's identity governance, right? And so, what I'm seeing a lot of organizations do and working with several of them, right, where they've slowed down a lot of their particularly agentic deployments and You know, we've, instead of like, you know, entitlements management, for instance, has always been an issue as long as, before cyber was even a thing, right? A dedicated field. And so, you know, so now with, you know, with a lot of these programs, we're baking in entitlements management or entitlements access reviews as AI enablement, right? And so we're doing cleanup projects, right? So how do you tackle this problem though? I mean, one of the ways that I tackle it is to turn something like entitlement management that's an abstract problem into something more concrete, right? Because I can go say all day long, fix your entitlement management, right? And like, what are the action steps to do that? If on the other hand I say, hey, the agent we plan to deploy would have the following permissions, right? I go create a persona account that has the following permissions and, you know, security groups, roles, um, and, and then I go and start looking for data, right? I use Copilot for M365, whatever tools are available to me with that persona to go look for data that this thing should absolutely not have access to, right? Every one of those, we've got to apply the Toyota 5 Whys method to it, right? So why, why did I have access as part of the security? Why am I part of the security group? Or why does the security group have access to this data? You follow it down to bedrock, right? Because each one of those is an exemplar of a, a typically a larger systemic problem, right? And once you find the root cause, you can go through and find, you know, other areas, right, where the same root cause, you know, was tripped, right? So, um, that's my approach.
Jonathan Knepher:
So I, I think this entitlement areas, actually, I kind of want to dig in a little more, right? Like, if I think about like the broad area, right, you've got, you've got like well-known agents or long-running agents and and so on, which makes sense. But you have individual developers, individual engineers, right? And they're using their own agents, even just in their IDE to do things.
Jake Williams:
Sure, yeah.
Jonathan Knepher:
How do you enable, like, those individuals then to control in a reliable way those entitlements? Like, does every engineer now have, like, a dozen entities of themselves that they have to individually manage? And how can they even do that?
Jake Williams:
So, I think this is where you've really got to get, you know, honestly, I think this is where you've got to get into governance. And so, you know, I apologize here, I'm looking for a Microsoft article I can point you to here. Yeah, here we go. So, I think we've got to get into identity governance first, right? And really step back and ask, I'm not talking about provisioning or anything, I'm literally talking about from a governance standpoint, can an agent run under a human identity? And you've got to ask that question. I have agents right now in my one-man LLC, right, that are running under my identity. In an enterprise, oh my gosh, that is nightmare fuel. It's nightmare fuel, right? So, you know, I think that's the first question I've got to ask, right? And the reason it's nightmare fuel, by the way, for the folks, you know, that might not otherwise, like, grasp this, you know, it's non-repudiation issues, right? And in fact, I just sent you a blog in the chat there too, right, that Microsoft published last year where they really walk through why OAuth, right, which is our— it's our go-to protocol for authorizing agents, has to evolve, right?
Jonathan Knepher:
Yeah.
Jake Williams:
You know, we still have OAuth servers, resource servers, right? I said OAuth servers, resource servers that are authenticated through OAuth. that still don't support on-behalf-of logging, right? And so, so then who did it, right? Is it the agent? Is it Jake? And then under what circumstances, right? And that's what the Microsoft— and even that Microsoft blog, they're, by the way, one of the founding members of the OIDC Foundation, right? The people that write the OAuth standards. So this is not like some random, you know, rant here. This is literally from, you know, yeah, literally from the mouth of the people that solved this. But I think that's what you gotta do first. You gotta determine, can I run agents under human identity or not? And then what are the rules of the road for that?
[24:39] CUSTODY and a Shared Language for Containment
Rachael Lyon:
Yeah. So, can we turn our attention to the framework that you built? I'd love for you to share a little bit more with our listeners. I believe it's called Custody. what it does and how it came about?
Jake Williams:
Yeah, so I've been in computer science a long time, one of the folks inside with a theoretical comp sci background. And my first experience with AI was agents, believe it or not, back in 2008, right? Agents are not new. What's new is the ability to run large language models and do natural language reasoning, right, with the agents. But agents themselves, you know, even non-deterministic agents aren't new. And so this problem of agentic containment was something that was pretty foreseeable. And I didn't expect— I was hoping to release this actually at a conference in October, and OpenAI and Anthropic said, hold my beer, and let a bunch of their agents out. And so I released it a little bit early with fewer worked examples than I'd hoped. But ultimately, Custody was basically an attempt to create both a taxonomy to discuss agentic behavior and then also have a schema, right, that we could use with CI/CD so that I could say, for an agent with these properties based on my governance standards, do these controls, the required controls, exist? And if not, break the build, right? Or at a bare minimum, alarm somebody.
I know break the build is a dirty word sometimes. Right? But, but at a bare minimum, you know, let, you know, basically let somebody, uh, let somebody know. But, but, you know, so far the biggest thing that I've seen come out of this, and I've got this adopted with a couple of clients already, um, you know, when I come in and I say, hey, what does this agent do? How much autonomy does the agent have? Um, is it expected to do tool calls, or is it just sending data to you, right? Or is it actually like modifying, you know, files and making API requests and sending data out of the network, right? So we're talking about something that's operational versus observational, excuse me, versus operational, or is adversarial, right? Is it one of these pen testing agents, right? Because, you know, in the first 2 cases, if it like changes identity, that's a bug. And the 3rd one, it could be a feature, right?
Jonathan Knepher:
Right.
Jake Williams:
It could be laterally moved by assuming someone else's identity. So as we look through even just defining those impacts of an agent, today, it takes me back and forth 20, 30 minutes with a competent engineer who actually understands the agent to get to brass tacks on, okay, what's the scope? And so, I've created a taxonomy of custody where we look at 3 areas. We look at the autonomy level, how autonomous is it? We look at the mandate of the agent. Observational, operational, adversarial, and then reach, right? Is it production data? Is it non-production data? If it is production, is it expected to have internet-wide access? And so now I can look and say, you know, my hope is down the road too, this will be more standard. We can look and say, this is an L3 observational R2 agent, and now we're all speaking the same language and it didn't take 30 minutes to get there, right? So yeah.
Jonathan Knepher:
How are you seeing the level of effort to implement this on those first couple of folks you're working with this on?
Jake Williams:
Yeah. So, I haven't had anybody yet go the full CI/CD route, right? There's a whole schema built out, right? Again, because I feel like if you don't, if something's only, you know, something that's happening on paper, right, like, you and I, I think, both know that we're not gonna get there, right? Or even to the effect that, you know, governance says go, we're still left to try to figure out, well, how do I actually implement this at machine level. So, so far, right, what we've done is in the 2 places that I have it implemented is that we've got the standard taxonomy done so that as somebody comes to governance and says, you know, hey, we are ready to deploy an agent, right, or, you know, IT infrastructure onboarding, we want to deploy an agent, we have governance standards set up around the different mandate, you know, autonomy levels and reach levels so that we can look and say, hey, if this is an L6 adversarial agent, No, right? That's just it. No. And so, so there's some red lines, and there's some others where we say, hey, here are the controls that are required. And so it's, it's really speeding up intake, you know, and getting like an idea already of just red line items, right? Where we say this combination can't be there, or if you're going to deploy this combination, here are the required security controls. Are you ready for those? No? Okay, then you're not ready to deploy. And so what we found then is just rapidly decreasing time spent in intake, you know, for these agents, right? That's been a big problem with a lot of my clients already, is they're getting— they're just drowning in security reviews for stuff that, from a risk perspective, is never going to be risk-approved, right? And so, this is cutting a lot of that time out already.
But we're not into the CI/CD builds yet. I can't wait to get there. That's where, you know, we ultimately want to go. But even if we just stop at the governance layer, you know, where we started applying a standardized taxonomy, that's still such a huge win for most organizations. So, yeah.
Rachael Lyon:
I'd like to come back to the autonomous aspect and kind of the conversations that you're having with clients on that front and how many are really kind of embracing fully autonomous and then the The part 2 of that, there's so much discussion on intent, right, which becomes the aspect of how do we measure that and, you know, to better get a handle on these things. But I'd be curious the conversations you're having around this.
Jake Williams:
Yeah. Yeah. So it's interesting because, you know, what I hear consistently, right, from CISOs is they'll say, well, the business wants autonomy, and I want control, right? And even some of the business units, once you get to the hey, you will by name be accountable for what this agent does, they also suddenly want more control.
Jonathan Knepher:
Right.
Jake Williams:
But what I tell folks is autonomy and control, they're opposite sides of the same balance scale, right? And so as you move autonomy up, you are necessarily moving control down. I mean, it's math. This is a little bit of mathematics, right? So what I'm hearing a lot is we want lots of control, right? But from a business enablement perspective, autonomy is what's giving them the ability to, you know, to be productive, right? So, you know, Jon even mentioned before, right, how do I review all those lines of code? Well, if you don't, that's autonomous, right?
Jonathan Knepher:
Right.
Jake Williams:
That was autonomously generated and committed and whatever without your human review code. So, you know, the best cases that I'm working with are folks that are really looking at not, hey, somebody brought me an agent and said, let's go deploy this agent. Right. One of my first questions is, what is your intended outcome? Mm-hmm.
Rachael Lyon:
Right.
Jake Williams:
Because I find that particularly— forgive me, this is going to sound gatekeepy, but I'm going to say it anyway. I'm finding a lot of people are suddenly engineers and are bringing me, you know, supposedly engineered solutions. And, you know, tell me what you want, right, and I'll help.
Rachael Lyon:
Right.
Jake Williams:
I'll help you build it securely, as opposed to, you know, a lot of the solutions we're seeing today or proposed today. And where I find we can be really effective with that, right, is in breaking down when somebody says, I want this agent to do all these things. You say, okay, well, you could do that, right, but that violates all kinds of least privilege. It has the lethal trifecta. There's literally no control around it. At best, I could put a monitoring control all in place. Or I can break your agent up into, in one case, 18 different agents, right? So planning agent, a supervisory agent, and then 16 task execution agents, each of which has very, very small, very, very small permissions, right? And then to talk agent to agent in a given workflow, they have to go through the supervisory agent. So now I have a pro-code gate, right, that's, that's, you know, that's being enforced, right? And so, now, I'm back to, is that move fast and break things? No, but it is move fast without— when I say move a little bit, enable you to move fast later, right? But we've gotta do the engineering upfront so that you don't break things, right? So, yeah.
Jonathan Knepher:
So, by the way, I've seen this scenario of everybody's an engineer now. And on one hand, it's beautiful, right? Like, you've got People building wonderful tools who aren't really coders.
Jake Williams:
Right.
Jonathan Knepher:
But yeah, it definitely makes me nervous.
Jake Williams:
Sure.
Jonathan Knepher:
Operationally, I love your idea of you have your supervisory agent, but how do you enforce that? How do you get— you've now got individual folks all around the company who've built their own little apps that are chugging away.
Rachael Lyon:
Yeah.
Jake Williams:
To put a note on that, right, one client I was working with earlier this year, they were like, yeah, we've got, you know, 2 dozen pilot agents in test, right? And I went into their Power Platform console and they had 2,400. And so, you know, it was only off by 2 orders of magnitude.
Jonathan Knepher:
Yeah, exactly.
Jake Williams:
You know, just a little bit, right? So I, you know, I hate to be the governance guy and come back and say like, look, either you've got to put governance in place, right, and standards in place. And again, this is going to introduce friction. It's going to slow things down. It's going to make them safer, or you accept the risks, right?
Jonathan Knepher:
Right.
Jake Williams:
And so, you know, that a citizen developer or whatever we're calling them, right? I'm just going to use Copilot Studio as the example here, right? They can go deploy a Copilot Studio agent. And so, in most cases, right, unless you've got solid governance, right, where you've locked down the connectors that they can use, right, you immediately have the lethal trifecta in every one of those agents. Right, or at least the capability for it, right? Lethal trifecta, of course, being, you know, access to sensitive data, an exfiltration route, and then access to untrusted data. And sometimes sensitive and untrusted data can come from the same data store, right? So, um, fun times, right?
Jonathan Knepher:
So I feel like you— like companies though need to then also invest in that common infrastructure, right? Rather than letting everybody run free, it's like you've got to give everybody the right sandbox to play in.
Jake Williams:
Sure. I mean, but isn't this SDLC all over again, right? You think SDLC, right? One of the things we start talking about, we're gonna standardize IDEs, we're gonna standardize the libraries that you use, the frameworks that you use, right? So that we have defined, you know, implementation patterns, right, that you can go back and building blocks that we're gonna use. And, you know, I mean, again, I feel like since we're talking about citizen developers, right? Being engineers, right? Why not apply SDLC principles there? And so, I agree with you 100%. We have to, like, it can't be everybody gets to pick your own framework, your own website that you like, agentic harness, whatever. It's standardized on a few because I can't govern everything, right?
Rachael Lyon:
So, I would love to come back kind of when you're talking to, like, CFO-level type folks and you know, kind of the assessing, you know, cost-benefit analysis or cost risks, and also looking at kind of the board level. I mean, basically what we're saying is business strategy, get comfortable with being uncomfortable. And I'm just wondering how those conversations are going for you.
Jake Williams:
Well, so just to level set here, you're talking about like the basically the secure I should say security, the ROI of an AI application. Is that what you're referencing?
Rachael Lyon:
Right, because obviously you're opening up a lot of risk. Oh, yeah.
Jake Williams:
Sorry. Gotcha. Yeah.
Rachael Lyon:
Yeah, exactly. How do you run a business on this questionable level of risk and cost and impact? But also versus we keep hearing a lot of competitive advantage and the people that get this really get it right. And do they become the leaders in an industry while others fall off? I mean, and it's such an interesting conversation. I'd be curious kind of the conversations that you're having in this realm.
Jake Williams:
Sure, sure. So, again, as I kind of referenced before, I try to take security out of it initially in the conversation because I think security is another— forgive me, I think it's an attribute in the overall cost and benefit type discussion. It's a consideration, but I don't know how to translate. I mean, the whole annual loss exposure thing, like, you're just finger in the air, kind of. I mean, that's such— come on, right? You know, I try to point back to first and foremost, before I get into the security of this thing, like, you know, does this actually make sense from a financial perspective, right?
Jonathan Knepher:
Mm-hmm.
Jake Williams:
And so, what we're seeing a lot of folks too that are not considering, like, they built an application, they look and based on some cost modeling, right, they say, oh, well, this is saving me X number of hours every week and it's costing a couple hundred bucks and whatever, and that works out to the, you know, works out. My first thing that I say is like, look, if you're bringing on, you know, if you can't tell me what you're saving with this thing, then what conversation are we even having, right? Because if you tell me you're saving a couple hundred dollars a week and I say you have, you've now introduced a risk of a regulatory violation, right, with data disclosure, right? That cannot possibly be risk— it just can't be risk justified. There's no way, right, for that level of, you know, savings. And that's really where I try to start the discussion, right, is to say first and foremost, like, you know, how close are you to like break-even here with ROI? And then we'll throw the risk attribute in and say, okay, for this benefit that you're getting, right, can you justify this risk? And that's where I try to get them down to it. I'm not trying to say this risk costs X dollars, right? It's to put the question because I look as a CISO, as a security engineer, I don't accept risks, right?
Jonathan Knepher:
Right.
Jake Williams:
I inform adults, right, who, thanks to our friends at Enron, there's a good, you know, literal chain, right, from the top all the way down. If you haven't been delegated to accept that risk, it's not your job anyway. So I try to put the information in front of the folks Right.
Rachael Lyon:
Yeah, that makes sense. Um, did you have another question, Jon? Because I'm, I'm segueing to my favorite part of the conversation.
Jonathan Knepher:
Yeah, go for it.
[39:18] Token Economics Nobody Wants to Hear
Jake Williams:
Real quick before you do that, right, I do want to throw one other thing in there, and that is that we are still not paying what tokens cost, right? 100%, right? And so that's something else that I— 2 other things that I point out to folks. We have the financial discussion. One, is, again, we're not paying what tokens cost. And 2, you know, so tokens will go up in cost over time. 2, if you want security monitoring, like the best defense against prompt injection today, right, is to take an input and before you send it to an LLM and say process this, you send it to a separate LLM session and do a classification task and say, does this look like prompt injection? Right? That is hella expensive, right? Regularly, I'm spending somewhere in the, you know, 15% to sometimes upwards of 30% on a given, you know, on a given workflow, you know, for that security. And that is not something that people are ready to hear because traditionally, right, we put a WAF in front of something, it's like it's a rounding error, right?
Rachael Lyon:
Yeah.
Jake Williams:
In the operational cost. It's not double, it's not a single-digit percentage, let alone double-digit percentages, right? So, those are the other things I try to bring into the costing discussion is, okay, you did this. Once I layer security on, are we still going to be okay? And once tokens cost what they cost, are we still going to be okay? And again, those are hard conversations to have, but they're necessary, right, in enterprise today.
Jonathan Knepher:
So, how far do you think the token cost is going to go, by the way? Like, in my own experience too, it's like I feel the tightening of the token cost, but yet the expansion of how many tokens are used, where are we gonna end up?
Jake Williams:
Yeah, well, that's such a great point, Jon, right? Because, you know, in so many cases, we're getting a new model out, right? And it is more capable. And then I go look at the same task I ran last week and, you know, used, I don't know, pick a number of tokens, right? I'm now using 1.2, 1.3x. And so, am I getting a marginally better result? Sometimes, right? But yeah, I— woof, man. I think there's a lot of space here for open weight models. I think a lot of organizations are going to start deploying their own, if also for nothing other than change control, right? We didn't even talk about change control. When you're running a model in, you know, running a model, you know, in Azure or using a Frontier model, you just have zero control, right? But if you're doing like Azure Foundry or Bedrock, AWS Bedrock, Google Vertex, whatever, they've got their own safety filters, AI safety filters in front of those that are changing on a regular basis with no notification to you, right? One of my own applications that I wrote for IONs broke because of this, right? So, CrowdStrike, obviously a very popular security firm, we have lots of tests involving prompts with CrowdStrike, and as we pushed it out to public beta among our clients, as opposed to the earlier closed beta. Folks are like, nothing's working here with CrowdStrike. We start looking, turns out Microsoft made a change to a safety filter, right? And it was interpreting CrowdStrike— I so wish I was making this up, right— as a call to violence.
Oh my heavens. Right? And so at that point, then it doesn't process the request. It does send a 200 series response code, right? Because from Microsoft's perspective, the API call was successful, right? So you made it to the— you authorized correctly, you made it to the authenticated, and you're authorized to make the call, but no, we're not answering, right? And so you can imagine some of the development headaches that come around for this, right? And so I think open weight models are going to be the future, both to control cost. You know, I have a one-time CapEx cost, you know, to get the GPUs in place, you know, to ultimately run the model. And it's fully predictable from there on out, right? I don't have something sitting in front of it that's doing the changes. And so I think some organizations are gonna move towards that. I honestly don't see how OpenAI or Anthropic, you know, survive, you know, trillion-plus dollar valuations. And I just— maybe I'm wrong, right? But I certainly would not be investing in either of those today, right? So, you know, I use the heck out of them.
I'm paying both of them. I wouldn't invest in either of them.
Rachael Lyon:
Right. I'll leave that there.
Jake Williams:
Not investment advice, right? There we go.
Rachael Lyon:
Not investment advice. So I always like to, you know, looking at time, I always love to wrap up our conversations with a little more kind of personal discussion. And I mean, I think first and foremost, our listeners are dying to know what's, what's with the t-shirt? What's going on there on that t-shirt?
Jake Williams:
Oh, so many questions. Yeah, Stop Building Hell. It's an anti-tech dystopia, right? I mean, as much as I work in the tech space and enable a lot of the companies that unfortunately are building the tech dystopia, we as a society should stop. And so that somebody put a— I don't remember the gentleman that put the shirt together. It was like a limited run, but one of the gentlemen, artists that I follow on Bluesky, and I was like, immediately, I'm like, yes, I want one of these. So here it is.
Rachael Lyon:
Yeah, I love it. Is that a beaver?
Jake Williams:
Yeah, it's a beaver. Yeah, they've got a whole like mantra at the bottom of the shirt, right?
Rachael Lyon:
Love it. Yeah, love it. That's fantastic. And then the final question, we're always curious on how did you get on the path to where you are today? It's not always linear. Some people start off and they're a PhD in music history and they're a CISO somewhere, or you know, they just kind of tinker on the side and then find themselves here, you know, later in life. But I would love if you could share your origin story of how you got here.
Jake Williams:
Woof. So, I got in trouble way back in the day, like, in my early teens, right? So, you know, 13, 14. My grandfather had been an engineer for GE. You know, he taught me binary and hex at, like, I don't even know, 6, 7, whatever. He's both an electrical and a mechanical. He's one of those folks that dual majored and whatnot. We ended up having— or he had a computer in his house, right? So I had a chance to go play with that and what have you. Young Jake did not know about signed numbers and signed versus unsigned numbers, right?
Rachael Lyon:
Mm-hmm.
Jonathan Knepher:
And prodigy.
Jake Williams:
Do you remember Prodigy back in the day?
Jonathan Knepher:
Oh, yeah.
Jake Williams:
Yeah. So, you had the dial-up service where you could get on the Prodigy network and you could send messages to people that were on Prodigy, but then they had this ARPANET thing that got connected, right? You had the email credits and nobody in my household, family, whatever, he was super stingy. I mean, good guy, right? But like, nobody's paying for those and I wanted them, right? I was like, I want these. And so, I figured out because it was all client-side, right, for did you have enough credits. And so I kept adding credits and, you know, rolled that number over to negative and broke some stuff in Prodigy. And anyway, so—
Jonathan Knepher:
And then got their attention, I assume, right?
Jake Williams:
Oh yes, yes, got their attention. Definitely got a lot of family attention. My, my mom then sent me down. She's like, done, you're off the computer, right? Mom's an Army nurse, you know, or I guess at the time, now reservist, right? Um, had a company command. And, and so she's like, you're going to the rescue squad, right? And she just, you know, booted me over there, this, uh, you know, small-town volunteer rescue and fire. Um, so I got my EMT at 16. Um, you know, and then, uh, uh, you know, I was in, uh, I was in pre-med. Um, and, uh, that did not work out well.
So, so didn't study well and ended up in the Army. The Army is like, you're not being a medic, right? And so they turned me around to intelligence and the rest is history. So, you know, one thing after another, just, you know, back, back to cyber roots. So, yeah.
Rachael Lyon:
Love it. Full circle, coming full circle. It was meant to be.
Jake Williams:
Crazy.
Rachael Lyon:
Wonderful. Well, thank you, Jake. This has been a wonderful conversation. Greatly appreciate your time and your insights, sharing that with our audience. And I will be sure to share those links in our show notes. Thank you so much for providing those as well. Yeah, definitely. Definitely.
To all of our listeners out there, uh, Jon, would you like a drum roll, please?
Jonathan Knepher:
Smash that subscribe button.
Rachael Lyon:
And you get a fresh new episode every single Tuesday. So until next time, everyone, stay secure.
About Our Guest

Jake Williams, VP of R&D at Hunter Strategy and a Faculty Analyst at IONS Research
Jake Williams, known in the security community as MalwareJake, is VP of R&D at Hunter Strategy and a faculty analyst at IANS Research, where he advises CISOs at organizations ranging from the Fortune 10 to mid-market. He has worked both sides of the keyboard. A former NSA-designated Master Network Exploitation Operator, he spent years in offensive operations and incident response before turning to advisory work, and he is a two-time winner of the DoD Cyber Crime Center's Digital Forensics Challenge. His current research sits where AI and machine learning collide with security, including CUSTODY, his open framework for classifying and containing autonomous agents. He also co-hosts the Breach Please podcast.