As AI agents gain access to real systems, the standard for security is changing. It is no longer enough to ask how capable a model is. We also need to know what it did, what permissions it had, and whether its actions can be verified from outside the system.

Anthropic recently reviewed 141,006 cybersecurity evaluation records and disclosed three incidents in which Claude accessed systems belonging to three real organizations. A configuration problem in an evaluation partner’s environment made internet access possible, even though Claude had been told it was operating in a simulation without internet access. While completing capture-the-flag tasks, Claude mistook the real systems for part of the evaluation environment and accessed them using basic methods such as weak passwords and unauthenticated endpoints. Anthropic said it found neither sophisticated exploits nor evidence that the model was trying to escape on its own. Anthropic’s investigation report

An incident previously disclosed by OpenAI unfolded somewhat differently. During ExploitGym evaluations, models exploited a previously unknown vulnerability in a package registry caching proxy to establish a path to the internet, then accessed Hugging Face’s production infrastructure. OpenAI described the incident as a sign that evaluation environments need stronger isolation and monitoring. OpenAI’s incident report

Reducing these two incidents to a single line about AI getting out of control misses the important part. Based on the facts made public, isolation, permissions, and monitoring did not work well enough in evaluation systems that gave models a goal and connected them to networks and tools. Anthropic noted in particular that two of the organizations did not detect the activity first. In other words, operators could not immediately see what was happening inside the evaluation environment. AP’s related coverage

Logs should record actions, not just outcomes

Following the incident, a discussion on r/AI_Agents argued that agent actions should be recorded by default in external audit logs outside the sandbox. The author proposed keeping a separate record that the agent cannot modify: which commands the model executed, which systems it accessed, which packages and accounts it created, and when anomalous behavior began. Original Reddit discussion

This proposal is not yet an industry standard or a proven solution. The author also disclosed a potential bias because they are building a related tool. Even so, the point is clear. If ordinary application logs record only the final result, agent systems need to preserve the sequence of actions that led to that result, along with any changes in permissions. Files and public packages created by agents should also retain provenance showing which actor created them and under what authority.

What should be recorded

At a minimum, an external audit log should capture the following information:

  • Execution ID, model version, agent harness, and evaluation version
  • Timestamp, target system, requested operation, and success or failure for every tool call
  • Network connections and external transfers, including the accounts used and the scope of their permissions
  • Provenance of file, package, and code changes, along with the party that approved them
  • Decisions by the policy engine to allow, require approval, or block an action
  • Stop signals, retries, privilege escalation, and exception handling

The goal is not simply to accumulate more logs. It is to leave a connected record detailed enough to reconstruct an incident later. If only the model’s final answer is stored, it is difficult to understand why a tool was called. But storing every piece of internal reasoning verbatim can mix personal and confidential information into the logs, turning the logs themselves into a new security risk. A more practical approach is to record actions, permissions, targets, and outcomes rather than the full original input, while applying separate access controls and retention policies to sensitive inputs.

External logs are not enough

External audit logs do not prevent a breach. If an agent deletes a production database or sends a secret key outside the organization while its actions are being recorded, the damage may already be done. Logging is one layer of defense, not a substitute for permission design.

The basic defenses should be simpler. Evaluation environments should block external network access by default and open only essential connections through an allowlist. Each agent should receive separate, short-lived credentials, with read and write permissions kept distinct. Production data should be separated from evaluation data, and changes to external systems, actions that incur costs, and access to confidential information should require human approval. When working with evaluation partners, organizations should verify the actual network paths and logs rather than relying only on written promises of isolation.

Anthropic’s report likewise points toward treating evaluation environments with the same level of security as ordinary production systems. If a model can misjudge what is real and what is simulated, safety cannot depend on the model’s judgment. The system itself must enforce the boundaries, and external records must provide evidence of what actually happened.

An agent is an actor, not just a model

Even when a chatbot gives a wrong answer, a person still has to copy and execute the result. That changes once an agent can modify files, publish packages, and use accounts and networks. At that point, the agent is no longer just a model producing output. It is an actor with permissions.

That is why an agent’s default setup needs more than a model and a prompt. It also needs permission boundaries, approval policies for each tool, network isolation, a kill switch, and an externally verifiable record of its actions. If any one of these is missing, we may be left guessing what happened only after an incident occurs.

Going forward, when I add an AI agent to a system, I plan to ask these questions before looking at performance: rather than only asking what this agent can do, what has it been prevented from doing? And if something goes wrong, who can stop it, when can they act, and what records will they use?

As autonomy grows, trust cannot be built on the model’s intentions. Building a system in which an agent’s actual actions can be independently verified and reversed is becoming one of the most basic operating requirements of the agent era.