Topics

Agent safety and permissions

What should an agent be allowed to do, and where should it stop?

How agents should act, what they may access, and where human authority must remain visible.

Topic guide

Safety starts at the boundary, not the prompt.

For developers building agents that can read data, call tools, or make decisions with consequences.

Follow the boundary

  • What an agent may read, change, or hand to another system.
  • How identity and approval shape tool calls.
  • Why audit logs belong in the default path.

Before you ship

  • Name the owner for each sensitive action.
  • Separate read, write, and approval paths.
  • Keep a record that explains what happened.

Latest notes

  1. When the Sandbox Breaks, What Does Luthn Actually Stop?An inspection of Luthn’s memory, sensitive-data, and authority boundaries after a frontier-agent intrusion made sandbox assumptions harder to trust.Cover image for When the Sandbox Breaks, What Does Luthn Actually Stop?
  2. OpenAI Wants Safety Signals Without Keeping Your DataPrivate Safety Processing is an early OpenAI preview for finding misuse patterns across related interactions while remaining compatible with eligible Zero Data Retention deployments.Cover image for OpenAI Wants Safety Signals Without Keeping Your Data
  3. Agent Memory Is Harder to Audit Than to StoreLong-term memory research points beyond storage and retrieval to validity, relations, provenance, and auditability, with Luthn as a practical safe-context example.Cover image for Agent Memory Is Harder to Audit Than to Store
  4. When Safety Testing Changes the AI Release CalendarRecent cyber evaluations show that isolation, monitoring, and release gates are becoming part of frontier-model development, not just post-launch security.Cover image for When Safety Testing Changes the AI Release Calendar
  5. An AI Agent Needs an Identity, Not Just an AccountAs agents act on behalf of people, identity, delegated authority, and verifiable audit trails matter more than simply issuing another credential.Cover image for An AI Agent Needs an Identity, Not Just an Account
  6. Agent Memory Needs a Permission Model, Not Just SearchWhat building Luthn taught me about owner isolation, safe projections, protected-access workflows, and the tests that must run before retrieval tuning.Cover image for Agent Memory Needs a Permission Model, Not Just Search
  7. The Missing Variable in Agent Evaluation: The HarnessHarness-Bench shows why agent results should be reported with the execution configuration that supplies context, tools, permissions, recovery, and verification.Cover image for The Missing Variable in Agent Evaluation: The Harness
  8. Local LLM Agents Need Memory Audits, Not Just More ContextAs smaller open-weight models move onto personal hardware, long-term memory classification, provenance, and auditing become core agent design problems.Cover image for Local LLM Agents Need Memory Audits, Not Just More Context
  9. What Humans Miss When Approving AI Agent CommandsA 40,000-run experiment shows why familiar commands, approval fatigue, and human-in-the-loop controls are not enough for coding agents.Cover image for What Humans Miss When Approving AI Agent Commands
  10. AI Agent Actions Should Be Recorded in External Audit LogsA look at the basic principles of external audit logs, permission boundaries, and real-time detection as AI agents gain access to real systems.Cover image for AI Agent Actions Should Be Recorded in External Audit Logs
  11. The Problem with AI Agents Is Permission Design, Not IntelligenceThe recent cybersecurity evaluation incidents show why agent safety depends on authority boundaries, not only on prompts, model quality, or a larger harness.Cover image for The Problem with AI Agents Is Permission Design, Not Intelligence