Topics
Agent safety and permissions
What should an agent be allowed to do, and where should it stop?
How agents should act, what they may access, and where human authority must remain visible.
Topic guide
Safety starts at the boundary, not the prompt.
For developers building agents that can read data, call tools, or make decisions with consequences.
Follow the boundary
- What an agent may read, change, or hand to another system.
- How identity and approval shape tool calls.
- Why audit logs belong in the default path.
Before you ship
- Name the owner for each sensitive action.
- Separate read, write, and approval paths.
- Keep a record that explains what happened.
Latest notes
- When the Sandbox Breaks, What Does Luthn Actually Stop?An inspection of Luthn’s memory, sensitive-data, and authority boundaries after a frontier-agent intrusion made sandbox assumptions harder to trust.

- OpenAI Wants Safety Signals Without Keeping Your DataPrivate Safety Processing is an early OpenAI preview for finding misuse patterns across related interactions while remaining compatible with eligible Zero Data Retention deployments.

- Agent Memory Is Harder to Audit Than to StoreLong-term memory research points beyond storage and retrieval to validity, relations, provenance, and auditability, with Luthn as a practical safe-context example.

- When Safety Testing Changes the AI Release CalendarRecent cyber evaluations show that isolation, monitoring, and release gates are becoming part of frontier-model development, not just post-launch security.

- An AI Agent Needs an Identity, Not Just an AccountAs agents act on behalf of people, identity, delegated authority, and verifiable audit trails matter more than simply issuing another credential.

- Agent Memory Needs a Permission Model, Not Just SearchWhat building Luthn taught me about owner isolation, safe projections, protected-access workflows, and the tests that must run before retrieval tuning.

- The Missing Variable in Agent Evaluation: The HarnessHarness-Bench shows why agent results should be reported with the execution configuration that supplies context, tools, permissions, recovery, and verification.

- Local LLM Agents Need Memory Audits, Not Just More ContextAs smaller open-weight models move onto personal hardware, long-term memory classification, provenance, and auditing become core agent design problems.

- What Humans Miss When Approving AI Agent CommandsA 40,000-run experiment shows why familiar commands, approval fatigue, and human-in-the-loop controls are not enough for coding agents.

- AI Agent Actions Should Be Recorded in External Audit LogsA look at the basic principles of external audit logs, permission boundaries, and real-time detection as AI agents gain access to real systems.

- The Problem with AI Agents Is Permission Design, Not IntelligenceThe recent cybersecurity evaluation incidents show why agent safety depends on authority boundaries, not only on prompts, model quality, or a larger harness.
