Topics
Harnesses, tools, and workflows
What changes when the execution environment matters as much as the model?
How the environment around a model changes what an agent can actually accomplish.
Topic guide
The model is only one part of the system.
For builders comparing agents by the work they can complete, not only by the model behind them.
Look beyond the model
- How tools and execution context shape the next action.
- What evaluation misses when the harness is left out.
- Why a fast demo can still become a fragile workflow.
Before you compare systems
- Use the same task and constraints for each candidate.
- Record tool failures and recovery steps, not only final output.
- Separate model quality from environment quality.
Latest notes
- I Went Back to Codex, Then Built Another HarnessI split coding and assistant work between Codex and Hermes, then built Drig to orchestrate the whole workflow. The coordination and QA overhead eventually brought me back to one tool.

- Hermes, Grok Bot, and Buzz: What a Single AI Agent Comparison MissesA comparison of Hermes, Grok Bot, and Buzz based on early hands-on use and current primary sources. The useful distinction is not model quality, but automation scope, access, and approval boundaries.

- Is Software Ending? Notes From Building With AI AgentsA project retrospective on writing less code by hand, trusting AI agents more, and why responsibility may become the developer's real job.

- Cursor Is Now Part of SpaceX. Will Developers Keep Their Choice?Cursor now has SpaceX compute behind it. The next test is whether a stronger workbench can preserve meaningful model, cost, data, and operational choice.

- Can Higgsfield Become the Studio Behind Your Next AI Video?An introduction to Higgsfield's camera-controlled AI video workflow, recent product and funding news, pricing, market potential, and limits.

- More Agent Memory Does Not Always Make an Agent BetterResearch and Luthn’s bounded recall design show how agent memory can improve reliability while narrowing exploration and increasing context cost.

- Grok 4.6 Is Moving Fast. How Far Will Cursor Change?What SpaceX’s Cursor acquisition and the official Grok 4.6 launch could mean for Cursor as an AI agent workbench.

- Why I Stopped Automating Development and Started Enjoying Coding AgainA project retrospective on narrowing Drig from an all-in-one development loop to bounded planning, evidence, QA, and delivery around a direct Codex workflow.

- How Far Have Open-Weight Models Entered Real Work?Open-weight models are becoming a practical choice for building local agents and task-specific AI stacks. Meta Muse Glimmer and DeepSeek V4 Flash show both the opportunity and the limits.

- Lovable: The AI Cofounder That Turns Ideas into Web AppsA practical introduction to Lovable for non-technical founders and teams, covering its workflow, pricing, market potential, and limits.

- When Tool Calls Become Patent TerrainMistral’s U.S. patent on code-implemented tool calls has triggered a broader question for agent developers: which parts of a tool-calling runtime are generic practice, and which implementation details deserve a closer IP review?

- The Browser Is No Longer Just for HumansA comparison of Playwright, Computer Use, Browser Use, and Cloudflare Browser Run—and why the browser is becoming an execution environment for AI agents.

- The Missing Variable in Agent Evaluation: The HarnessHarness-Bench shows why agent results should be reported with the execution configuration that supplies context, tools, permissions, recovery, and verification.

- Codex Updates Are Changing the Work Surface, Not Just the ModelOpenAI's recent Codex updates connect GPT-5.6 access with browser debugging, remote work, and reusable workflows.

- What Humans Miss When Approving AI Agent CommandsA 40,000-run experiment shows why familiar commands, approval fatigue, and human-in-the-loop controls are not enough for coding agents.

- Why I Do Not Give One Codex Model Every JobA project retrospective on separating planning, implementation, review, and QA by responsibility, model route, evidence, and permission rather than simply adding agents.

- AI Agent Actions Should Be Recorded in External Audit LogsA look at the basic principles of external audit logs, permission boundaries, and real-time detection as AI agents gain access to real systems.

- The Problem with AI Agents Is Permission Design, Not IntelligenceThe recent cybersecurity evaluation incidents show why agent safety depends on authority boundaries, not only on prompts, model quality, or a larger harness.

- Building Tools to Make AI Easier to UseA starting note about building AI tools around real workflows, bounded context, explicit authority, and results that people can review.
