
Teams almost always try to build Zero Trust for AI agents from the wrong end: access policies, firewalls, detection. The place to start is duller and more important: how many agents do you have, and who launched them. You cannot govern what you cannot see, and in nearly every company we work with, the real list of autonomous AI systems diverges sharply from what the security team believes is running.
The METR incident is a good example. METR is a nonprofit that tests AI models. An attacker found a personal EC2 instance running an agentic application, bypassed authentication, and over three weeks drained 600,000 dollars in tokens through a stolen API key. The internal monitoring dashboard did not show the rate-limited requests, so the anomaly never landed in front of the people who were supposed to catch it. The problem here is not broken access control. The problem is that the instance was never in the inventory of systems anyone was watching.

Three blind spots that make agent Zero Trust fail
We run AI agent penetration tests regularly and see the same pattern every time: organizations jump straight to controls and skip the visibility step. The cause is almost always one of three.
Agents as shadow IT
Agent adoption outruns agent governance by months, sometimes years. A developer wires an agent SDK into an internal service on a Friday evening to close out a sprint. By Monday that agent has a standing access token to the production database that nobody but its author knows about.
Security’s first instinct is to block anything unfamiliar. But blocking before you have an inventory does not remove the risk. It pushes people around the policy, and the agents go further into the shadows instead of disappearing. The approach that works is to treat spend on AI tooling as a signal to inventory, not as a violation. A new token provider shows up in cloud billing or API logs, open a record for it, not an investigation. And give people an official fast path to get an approved tool, or the workaround stays, it just gets quieter.
Blind spots across layers of the stack
No single telemetry source sees the whole agent fleet. An agent can sit in the network, on an endpoint, in the browser as an extension, inside a SaaS integration. TLS encryption hides the prompts and tool calls themselves from network monitoring.
Typical example: a browser extension summarizes customer data from the CRM and drafts emails. Endpoint EDR does not see it, and the traffic to the SaaS domain is encrypted. To an attacker it is a ready-made foothold: a compromised extension pulls data out of the CRM and none of the usual controls notice.
This does not close with one tool. It closes with correlation across disparate signals: DNS and SNI queries, JA4 fingerprints of TLS connections, host environment variables, granted OAuth scopes, browser telemetry itself. Any one signal is weak on its own. Together they add up to a real picture of what your agents are doing on your network.
The gap between the audit and reality
An annual or even quarterly audit is stale the moment the report ships. Agentic systems clone, deploy, and get torn down in minutes. A short-lived agent clone can spin up, exfiltrate data, and vanish faster than the interval between two scheduled reviews, and an attacker who knows a given organization’s audit cycle paces the attack to the gaps between them.
There is one answer that works: continuous monitoring instead of a point-in-time snapshot. In practice that is a system watching agents constantly, including an agent watching other agents; regular review of granted permissions rather than a once-a-year reconciliation; a named owner on every agent; and a working kill switch, an emergency shutoff you can actually hit mid-incident, not just describe in a policy document.

The order of operations that actually works
The right sequence is simple: inventory and governance first, then protective controls, and only then detection. Not the other way around. In practice it breaks down into concrete steps:
- Find and inventory every agent. Not “roughly aware of,” but an actual list with owners, each agent’s purpose, and what it has access to.
- Give each agent its own identity. An agent should not inherit the full permission set of the user who deployed it by default. This is one of the most common findings in our infrastructure security assessments.
- Set up continuous monitoring instead of quarterly one-off checks.
- Put an authorization layer between the agent and the services it reaches, instead of a straight pass through inherited permissions.
- Log the tool calls, not just the prompts. What an agent actually did shows up in the calls, not in the text of the request.
- Only after that, deploy blocking controls, and only against a known, inventoried set of agents, not blindly against everything.
Where our own work fits
We work this problem from several angles across our own writing and products, and it keeps reducing to one thing: an agent’s identity and visibility matter more than trying to predict its behavior. We covered this in detail in our piece on secrets sprawl as a non-human identity problem.
For inventory and detection of agents on the perimeter, including someone else’s, we have BastionAgent. Instead of trying to block an unknown agent in advance, it seeds dead watermarked credentials and traps that only a reasoning system would engage with. The act of an agent interacting with your perimeter becomes a logged event. We walked through a similar scenario with the research agent that climbed over a government portal’s access control. The threat model is the same: an agent that should never have gotten there became visible only at the moment it acted.
For your own agents that can go rogue on the outbound side, our firewall airlock_ai stops an agent from exfiltrating data or using credentials it should never have touched, whether or not you predicted that behavior in advance.
The discovery and assessment phase itself is part of a broader offensive discipline. We have written about why a single model does not cover even half the picture, and our Brain was built around that idea: route tasks across different models and merge findings into one picture rather than leaning on a single source of truth. If a leak or foothold has already happened, the work moves to an assumed-breach assessment, and building this discipline into CI/CD and the agent development pipelines from the start is what secure code review is for.
The takeaway
Zero Trust for AI agents is not a set of policies you configure once and forget. It is a discipline, and it starts with the boring, unfashionable question “how many agents do we have and who owns them,” and moves to protective controls and detection only after that. Organizations that try to leap straight to blocking and alerting repeat the METR story: the control may have been in place, but nobody was looking where the leak was actually happening. Start with the inventory, give every agent its own identity and owner, set up continuous rather than point-in-time control, and only then add blocks and automated response.