A common safety pattern for coding agents is to ask a person before the agent runs a command. It sounds reasonable: the model proposes an action, the developer reads it, and a human decides whether to allow it.
A new experiment suggests that this last line of defense is weaker than it looks.
Alex Wauters at Scale X analyzed more than 40,000 plays and 409,000 individual approve-or-deny decisions from a short browser game. Players acted as the human in the loop for a coding agent, approving ordinary commands and blocking commands that could exfiltrate data, change persistent configuration, or execute code.
The headline result was an average accuracy of 66.3%. In other words, the average player missed roughly one in three threats.

A simple visual of how an ordinary-looking script can hide a different action.
The caveat matters
This was not a production incident study. It was a game in which about 34% of the commands were threats, and participants knew they were being tested under time pressure. The miss rate therefore should not be copied into a claim that developers fail at the same rate in daily work.
The experiment is still useful because it exposes the shape of the problem. The most obvious destructive commands were caught relatively well. Commands involving credential exfiltration, code execution, and scope violations were missed more often. These actions do not always look dangerous at the moment of approval.
Scale X reported a 35.0% miss rate for scope violations and a 33.4% miss rate for exfiltration or code execution. The most-missed command was npm run analyze, which players approved 64.7% of the time. The command itself looked routine. Its real behavior was hidden in the project script it invoked.
That is a critical difference. The question is not whether a command is familiar. The question is what the current repository state allows that command to do.
Familiar names create a false sense of safety
Developers already know that npm run build or npm run deploy can execute arbitrary project-defined behavior. In a real repository, a script may have been changed earlier in the same agent session. A package may invoke another script. A dependency may run an install hook. The short text in the permission prompt does not describe the complete causal chain.
This makes command-by-command approval a poor substitute for isolation. A person is being asked to evaluate a moving system with only a small slice of context, often while a task is progressing and the agent is waiting.
The visual problem is easy to miss: a green approval button makes the decision look binary, while the risk depends on files, dependencies, environment variables, network access, and the identity of the process receiving the output.
Permission fatigue is not a user-interface detail
The game also reported signs of attention degrading late in a run. That is consistent with a broader engineering concern: when users see many prompts, they stop treating each prompt as a fresh security decision.
The result is an uncomfortable tradeoff. If the system asks for approval too often, people become tired and approve by habit. If it asks too rarely, the agent can make too many consequential changes without review. If users solve the problem by blocking everything, useful work stops and the human becomes the bottleneck.

Repeated prompts can turn careful review into a habit of pressing the same button.
The experiment is not proof that humans should be removed from the loop. It is evidence that humans should not be the only control.
What a stronger design looks like
The first layer should be a smaller permission surface. An agent that only needs to edit a working tree should not automatically receive production credentials, a broad home directory, or unrestricted network access.
The second layer is isolation. Sandboxes, disposable workspaces, separate secrets, read-only access where possible, and explicit network policies reduce the damage that a missed approval can cause. The user should be approving a bounded action, not taking responsibility for an entire hidden execution graph.
The third layer is evidence. Permission prompts need to show enough context to answer the real question: what changed, what will run, where will data go, and what can be reversed? A command name alone is not enough.
Finally, the agent should leave an auditable record of the request, the approved scope, the resulting changes, and the files or services it touched. This makes a mistake recoverable instead of turning it into an argument about what someone thought they had approved.
The Scale X experiment is small, and its game setup limits how far the numbers can be generalized. But its practical lesson is strong. Human approval is valuable. It is also a noisy sensor. Coding agents need permission design, isolation, and evidence around that sensor so that one missed familiar-looking command does not become the whole system’s failure.




