Topics

Harnesses, tools, and workflows

What changes when the execution environment matters as much as the model?

How the environment around a model changes what an agent can actually accomplish.

Topic guide

The model is only one part of the system.

For builders comparing agents by the work they can complete, not only by the model behind them.

Look beyond the model

  • How tools and execution context shape the next action.
  • What evaluation misses when the harness is left out.
  • Why a fast demo can still become a fragile workflow.

Before you compare systems

  • Use the same task and constraints for each candidate.
  • Record tool failures and recovery steps, not only final output.
  • Separate model quality from environment quality.

Latest notes

  1. I Went Back to Codex, Then Built Another HarnessI split coding and assistant work between Codex and Hermes, then built Drig to orchestrate the whole workflow. The coordination and QA overhead eventually brought me back to one tool.Cover image for I Went Back to Codex, Then Built Another Harness
  2. Hermes, Grok Bot, and Buzz: What a Single AI Agent Comparison MissesA comparison of Hermes, Grok Bot, and Buzz based on early hands-on use and current primary sources. The useful distinction is not model quality, but automation scope, access, and approval boundaries.Cover image for Hermes, Grok Bot, and Buzz: What a Single AI Agent Comparison Misses
  3. Is Software Ending? Notes From Building With AI AgentsA project retrospective on writing less code by hand, trusting AI agents more, and why responsibility may become the developer's real job.Cover image for Is Software Ending? Notes From Building With AI Agents
  4. Cursor Is Now Part of SpaceX. Will Developers Keep Their Choice?Cursor now has SpaceX compute behind it. The next test is whether a stronger workbench can preserve meaningful model, cost, data, and operational choice.Cover image for Cursor Is Now Part of SpaceX. Will Developers Keep Their Choice?
  5. Can Higgsfield Become the Studio Behind Your Next AI Video?An introduction to Higgsfield's camera-controlled AI video workflow, recent product and funding news, pricing, market potential, and limits.Cover image for Can Higgsfield Become the Studio Behind Your Next AI Video?
  6. More Agent Memory Does Not Always Make an Agent BetterResearch and Luthn’s bounded recall design show how agent memory can improve reliability while narrowing exploration and increasing context cost.Cover image for More Agent Memory Does Not Always Make an Agent Better
  7. Grok 4.6 Is Moving Fast. How Far Will Cursor Change?What SpaceX’s Cursor acquisition and the official Grok 4.6 launch could mean for Cursor as an AI agent workbench.Cover image for Grok 4.6 Is Moving Fast. How Far Will Cursor Change?
  8. Why I Stopped Automating Development and Started Enjoying Coding AgainA project retrospective on narrowing Drig from an all-in-one development loop to bounded planning, evidence, QA, and delivery around a direct Codex workflow.Cover image for Why I Stopped Automating Development and Started Enjoying Coding Again
  9. How Far Have Open-Weight Models Entered Real Work?Open-weight models are becoming a practical choice for building local agents and task-specific AI stacks. Meta Muse Glimmer and DeepSeek V4 Flash show both the opportunity and the limits.Cover image for How Far Have Open-Weight Models Entered Real Work?
  10. Lovable: The AI Cofounder That Turns Ideas into Web AppsA practical introduction to Lovable for non-technical founders and teams, covering its workflow, pricing, market potential, and limits.Cover image for Lovable: The AI Cofounder That Turns Ideas into Web Apps
  11. When Tool Calls Become Patent TerrainMistral’s U.S. patent on code-implemented tool calls has triggered a broader question for agent developers: which parts of a tool-calling runtime are generic practice, and which implementation details deserve a closer IP review?Cover image for When Tool Calls Become Patent Terrain
  12. The Browser Is No Longer Just for HumansA comparison of Playwright, Computer Use, Browser Use, and Cloudflare Browser Run—and why the browser is becoming an execution environment for AI agents.Cover image for The Browser Is No Longer Just for Humans
  13. The Missing Variable in Agent Evaluation: The HarnessHarness-Bench shows why agent results should be reported with the execution configuration that supplies context, tools, permissions, recovery, and verification.Cover image for The Missing Variable in Agent Evaluation: The Harness
  14. Codex Updates Are Changing the Work Surface, Not Just the ModelOpenAI's recent Codex updates connect GPT-5.6 access with browser debugging, remote work, and reusable workflows.Cover image for Codex Updates Are Changing the Work Surface, Not Just the Model
  15. What Humans Miss When Approving AI Agent CommandsA 40,000-run experiment shows why familiar commands, approval fatigue, and human-in-the-loop controls are not enough for coding agents.Cover image for What Humans Miss When Approving AI Agent Commands
  16. Why I Do Not Give One Codex Model Every JobA project retrospective on separating planning, implementation, review, and QA by responsibility, model route, evidence, and permission rather than simply adding agents.Cover image for Why I Do Not Give One Codex Model Every Job
  17. AI Agent Actions Should Be Recorded in External Audit LogsA look at the basic principles of external audit logs, permission boundaries, and real-time detection as AI agents gain access to real systems.Cover image for AI Agent Actions Should Be Recorded in External Audit Logs
  18. The Problem with AI Agents Is Permission Design, Not IntelligenceThe recent cybersecurity evaluation incidents show why agent safety depends on authority boundaries, not only on prompts, model quality, or a larger harness.Cover image for The Problem with AI Agents Is Permission Design, Not Intelligence
  19. Building Tools to Make AI Easier to UseA starting note about building AI tools around real workflows, bounded context, explicit authority, and results that people can review.Cover image for Building Tools to Make AI Easier to Use