Skip to main content

Your AI Risk Management Program Can't See Its Agents

|

0 分の読み物

See how Forcepoint stops AI data risk
  • Lionel Menchaca

Most AI risk management programs read well in a governance deck. They cite the right framework, name an accountable executive and include a risk register with the appropriate columns. Then someone asks a simple question: which agents touched customer data this week, and what did they do with it? For most security teams, there is no confident answer. That gap, between the policy an organization can produce and the visibility it actually has, is where AI risk management programs stall.

It stalls in a specific, predictable place. Governance policy is not the hard part. Neither is picking a framework. The hard part is Map and Measure, the two functions of the NIST AI Risk Management Framework that require an accurate, current picture of what data is actually flowing into and out of AI systems. Without that picture, an organization cannot map risk to a real system, and it cannot measure exposure it cannot see. This is not a governance failure. It is a visibility failure wearing a governance label.

What AI Risk Management Actually Covers

AI risk management is the ongoing discipline of identifying, evaluating and controlling the risks an organization takes on by building, buying or using AI systems. It spans security risk, such as data exfiltration through prompts or compromised agents, along with compliance risk, operational risk and the reputational risk that follows a public failure. It is a program, not a project, and it runs for as long as the organization uses AI.

An AI risk assessment is one exercise inside that program, not the program itself. It is the point-in-time or periodic review that inventories AI systems, classifies their risk level and documents where controls exist and where they do not. A good AI risk assessment feeds the broader AI risk management program with current evidence. A stale one gives a false sense of coverage, which is arguably worse than no assessment at all, since it produces confidence without accuracy.

For security practitioners, the distinction matters because most published guidance conflates the two. Treating an annual AI risk assessment as equivalent to an AI risk management program is how organizations end up with a report that accurately described their AI footprint eight months ago and says nothing useful about the three new copilots and one autonomous agent that went live since.

The NIST AI Risk Management Framework, in Practice

Security teams evaluating an AI risk management framework tend to land on the NIST AI RMF, and for good reason. It is vendor neutral, referenced in EU AI Act guidance, and built around four functions that translate cleanly into program structure: Govern, Map, Measure and Manage.

Govern sets the foundation

Govern is the cross-cutting function that establishes policy, accountability and organizational culture around AI risk. It answers who owns AI risk decisions, what the acceptable use policy allows and how third-party AI tools get evaluated before adoption. Most organizations have made real progress here. Governance committees exist. Acceptable use policies get published. This is the function most audits find in reasonable shape. The failure shows up one step later: policy gets written and published before anyone actually maps what exists to govern, the discovery-before-policy problem that undermines governance programs generally. It resurfaces immediately in Map, which depends on the discovery that Govern assumed already happened.

Map establishes context

Map requires understanding the specific context of each AI system: what it does, what data it touches, who depends on it and what could go wrong. This is where the framework starts asking questions that governance policy alone cannot answer. A policy can state that sensitive data should not enter unsanctioned AI tools. Map requires knowing whether it already has, and through which tool.

Measure tests against reality

Measure evaluates AI systems against the trustworthiness characteristics the framework defines, using both qualitative and quantitative methods. It is the function that turns a mapped risk into a scored one. Measure depends entirely on Map being accurate. An organization cannot measure exposure in a system it has not correctly mapped, and it cannot correctly map a system it cannot see in the first place.

Manage responds to what Measure finds

Manage prioritizes and responds to the risks Measure has scored, through technical controls, procedural safeguards or a decision to accept a given risk. Manage is only as good as the Measure output feeding it. A management decision built on an incomplete measurement is a decision built on a blind spot, no matter how well documented the response plan looks.

Where Map and Measure Break Down

The August 2025 OpenText and Ponemon Institute survey of nearly 1,900 CIOs, CISOs and IT leaders found that 53% describe reducing AI security and legal risk as very or extremely difficult, even as most treat AI adoption as a top priority. That gap between priority and execution tracks almost exactly with where Map and Measure fail in practice.

Three patterns show up repeatedly. First, shadow AI. Employees adopt AI tools directly, through browser extensions, free-tier accounts and embedded features in software the organization already licenses, well ahead of any sanctioned rollout. A risk assessment built on the approved vendor list is mapping a fraction of what is actually running. Second, autonomous agents. Agents call APIs, read and write data across enterprise systems and take multi-step actions without a human reviewing each step. Traditional monitoring built for human users and static applications was not built to attribute an action to an agent, correlate it with the data involved or flag it as anomalous in real time. Third, inconsistent classification. Data that is correctly classified in a file share often loses that classification the moment it moves into a prompt, a training set or an agent's working context, which means the sensitivity travels but the control does not.

Each of these breaks Map before Measure even gets a chance to run. An assessment that misses shadow AI has an incomplete inventory. An assessment that cannot attribute agent actions has no way to measure agentic risk with any confidence. An assessment built on data whose classification does not survive the move into AI is measuring the wrong thing entirely.

The Data Layer Is the Missing Instrumentation

The instinct in response to these gaps is usually to add more AI governance: more policy, more documentation, more committee review. That instinct addresses Govern, which was rarely the weak function to begin with. It does not address Map or Measure, because those functions are not governance problems. They are visibility and enforcement problems that sit at the data layer.

An AI risk management program can only map what it can see and measure what it can classify continuously, not just at the moment data was created. That requires visibility into shadow AI usage across the organization, real-time inspection of prompts reaching sanctioned and unsanctioned tools alike, and an audit trail for autonomous agent actions that can attribute what an agent did back to the data and the person or process that authorized it. Without that instrumentation, Map produces a snapshot that is out of date the day it is published, and Measure produces a score built on incomplete inputs.

This is a different starting point than most AI risk management guidance offers, and it changes the order of operations. Rather than starting with a governance framework and hoping the organization's existing tools can supply evidence for it, the data layer becomes the instrumentation the framework runs on. Map and Measure stop being manual documentation exercises and become outputs of continuous, automated visibility.

What This Looks Like in Practice

In practice, closing this gap means a few concrete capabilities working together rather than a single new governance process layered on top of existing tools. Continuous discovery of AI usage, sanctioned and unsanctioned, gives Map a current inventory instead of a point-in-time survey. Real-time inspection of prompts and outputs, correlated with existing data classification, keeps sensitive data from losing its protection the moment it enters an AI workflow. An audit trail that attributes agent actions to the identity behind them, human or automated, gives Measure something concrete to score instead of an assumption about agent behavior.

Forcepoint AI Data Security is built around this exact problem: protection that starts at the data layer and follows sensitive information into every prompt, agent and sanctioned or unsanctioned AI application it reaches. Rather than treating AI governance and data security as separate initiatives that each produce their own partial evidence, the two run on the same policy and the same classification, which means the evidence Map and Measure need is a byproduct of normal operation rather than a quarterly scramble to reconstruct it.

The Framework Was Never the Hard Part

The NIST AI Risk Management Framework is not the obstacle standing between most organizations and a defensible AI risk management program. Govern, Map, Measure and Manage are a sound structure. Most organizations that struggle with AI risk management are not struggling because they picked the wrong framework, and the fix is not another round of policy writing. They are struggling because Map and Measure, specifically, demand a level of data visibility their current tools were never built to provide. That is a narrower and more technical problem than a governance program that shipped before discovery happened. It is what is left over after governance gets fixed.

Closing that gap is the practical work of Self-Aware Data Security: knowing where sensitive data is the moment it is created, adapting protection as it moves into AI workflows, and giving security teams the evidence to prove it rather than describe it. That is what turns an AI risk management framework from a document into a program that actually holds up when a regulator, a board or an auditor asks for proof.

See how Forcepoint AI Data Security gives security teams the visibility and enforcement layer their AI risk management program is missing.

  • lionel_-_social_pic.jpg

    Lionel Menchaca

    Lionel Menchaca has covered data security at Forcepoint since 2020, writing about DLP, DSPM, insider risk and AI security for security and IT leaders. He works with Forcepoint X-Labs threat researchers to turn their findings on emerging threats, from AI-targeted supply chain attacks to prompt injection, into practical guidance, and he leads the company's editorial strategy across the blog and the X-Labs newsletter. Before Forcepoint, Lionel founded and ran Dell's corporate blog for seven years and spent two decades helping enterprise tech companies explain security, cloud and AI.  

    の記事をもっと読む Lionel Menchaca

X-Labs

インサイトや分析、ニュースを直接お届けします

要点

サイバーセキュリティ

サイバーセキュリティの最新トレンドや話題をカバーするポッドキャスト

今すぐ聴く