The recent OpenAI and Anthropic cybersecurity evaluation incidents are easy to describe as stories about AI getting out of control. The more useful question is narrower: what authority did the evaluation environment give the model, and which boundary failed to hold?
The systems gave models a goal and connected them to the internet, accounts, and tools that could help achieve it. The public reports describe what happened in those evaluations; they do not prove that every agent behaves this way in ordinary production. They do show why a prompt is not a permission system.
Writing “do not access the internet” in a prompt is entirely different from blocking network access. Agent safety is determined by the execution environment as much as by the model: network policy, identity, tool permissions, logs, approval gates, and a way to stop the run.
The incident pattern is an authority problem
An agent needs at least three things to act: a goal, a capability, and a boundary. A goal says what it is trying to achieve. A capability says which systems and tools it can reach. A boundary says what it may not do, when it must ask, and how the run ends.
| Layer | The question to answer | Typical failure |
|---|---|---|
| Goal | What outcome is the agent responsible for? | A broad objective is treated as permission to improvise. |
| Capability | Which accounts, files, network routes, and tools are available? | A model can reach more systems than the task requires. |
| Boundary | What needs approval, time limits, or a hard stop? | A warning in the prompt is expected to stop a real action. |
| Evidence | What records prove what happened? | A final answer hides the tool calls and state changes behind it. |

The public evaluation reports make the execution environment part of the safety discussion. Source: Anthropic
That is the practical lesson from the reports. The important fix is not necessarily a more intelligent model. It may be a smaller network scope, a separate identity, a read-only tool, an approval step, or a validator that refuses an incomplete artifact.
A more complex harness is not always better
I noticed a related problem while applying the harness and loop-engineering approaches that had become popular. Each part sounded reasonable: make a plan, divide the work, delegate it to multiple agents, review the result, retry failures, and summarize intermediate output. But when I added these capabilities one by one, the process became larger than the work itself.
At times, it felt like a hungry hippo devouring time and tokens.

The AP coverage is useful context for the incident discussion, but a news report is not a replacement for the primary evaluation record. Source: AP News
I have since stripped away most of that process and am experimenting with a simpler harness suited to my own work. That is a design judgment, not a controlled productivity result. A smaller harness is not automatically safer either; if it exposes a broad credential or lacks a stop condition, it can be simpler to operate and easier to misuse.
What a small harness must still make explicit
Before adding another planner or reviewer, I would check the boundary itself:
- Is each tool limited to the files, services, and network routes this task needs?
- Does the agent use a separate identity whose authority can be revoked?
- Are external changes and irreversible actions held behind a human approval?
- Are time, retry, and token budgets visible rather than implicit?
- Can a reviewer see the request, tool calls, result, and changed artifact?
- Does a failure return a bounded error instead of widening access?
This is why I now see agent design as closer to permission design than to organization design. Human teams need roles because people coordinate across a large social system. A small agent task needs only the roles that create a real control or evidence boundary. Adding a role because a human organization has one does not create safety by itself.
A concrete permission sequence
When applying an authority boundary to real work, it is easier to explain the system if the agent does not receive the whole job at once. A repository bug fix, for example, can be split into stages:
| Stage | Authority granted | Evidence to retain |
|---|---|---|
| Inspect | Read relevant files and run tests | Files read, commands run, test results |
| Propose | Write a change plan and draft patch | Files, rationale, and expected impact |
| Approve | Apply only the change a person accepted | Approver, time, and approved scope |
| Verify | Run checks in the bounded environment | Logs, artifacts, and failure reason |
| Close | Revoke credentials and remove temporary files | Revocation and final state |
In this sequence, the agent may read first but cannot immediately write to an external system. Network access can stay disabled when it is not needed. A deployment or data change can require a separate approval token tied to the task, time window, and target. A token that means “allow everything for this session” quietly removes the boundary the design was meant to create.
Failure handling is part of the permission design too. When a test fails, the agent should return the failure and the question a person needs to answer instead of automatically opening more files or services. Retries may increase, but the authority should not expand silently. This is a design example for making permissions reviewable, not a universal implementation or a measured productivity result.
What this article does not prove
The cited incidents were evaluations with their own goals, tools, and conditions. They should not be copied into a claim that developers or models fail at the same rate in daily work. Nor does my simplified harness prove that orchestration is bad. The supported conclusion is more practical: define the authority before optimizing the process.
Small goals, limited permissions, short feedback loops, and clear stopping points are a better starting point than a workflow that grows until nobody can explain why it is safe. As agents become more capable, the question is not only what they can do. It is which actions they were authorized to do, and what evidence remains when they stop.
OpenAI security incident report · Anthropic evaluation incident investigation




