I did not stop using automation. I stopped asking one graph to own the entire development loop.

For a while, I was building Drig around harnesses, graphs, and multiple agents. The original idea was reasonable: split development into planning, interviewing, implementation, review, and delivery, then connect those stages into a repeatable loop. Each agent would have a role, a bounded responsibility, and a defined handoff.

That design came from a real concern. Coding agents can lose context, skip a check, or make a decision that is difficult to reconstruct later. A structured harness seemed like a way to make development more visible and repeatable. I was not trying to automate coding because coding was unimportant. I was trying to protect the quality of coding by making the process more disciplined.

A busy set of workflow tubes, checklists, timers, and agent nodes surrounding a small code task

The more stages I added, the more effort went into moving work through the system.

The original bargain

The bargain looked simple on a diagram:

Stage What I wanted What the stage also introduced
Planning and interview Fewer ambiguous requirements More context to compile and hand off
Implementation A bounded coding responsibility Another boundary to explain and recover
Review and QA Independent checks More artifacts, timing, and state transitions
Delivery and closeout A reproducible final path A graph that could become the product being maintained

Over time, the system grew beyond the problem it was supposed to solve. I spent more time maintaining transitions, describing agent responsibilities, and recovering from automation edge cases. When a test failed, the first question was no longer only what was wrong with the code. I also had to find which stage produced the result, what context crossed the handoff, and whether the next agent interpreted it correctly.

The cost was not one dramatic bug. Every extra stage consumed tokens. Every handoff created another place where context could be shortened, distorted, or repeated. Every recovery path required design and testing. The system was meant to make development faster, but I was increasingly developing the system instead of the product.

There was a visibility problem as well. A lot was happening inside the automation, but it was not always obvious what had been generated, which agent had made a decision, or why the next stage had been selected. The more automatic the loop became, the more I felt like an operator watching a machine rather than a developer making a change.

The reset was a change of boundary, not a rejection of rigor

I eventually decided to drop Drig as an all-in-one development automation project. That did not mean throwing away tests, review, or evidence. It meant moving the boundary between direct development and orchestration.

The simplification plan at the time put product implementation in a direct Codex workflow and left planning, post-development checks, and delivery to Drig. Later documentation included reviews for each checkpoint again. The scope of review differs between those two versions, but both separate writing product code from deciding whether the work is complete.

I revisited the project documentation on September 5, 2026, and summarized the flow below. Planning defines a small piece of work and its acceptance criteria. During implementation, checkpoints preserve changes and validation evidence for review and QA. Only changes that pass QA proceed to a pull request. Merging remains an explicit owner decision.

Drig planning, direct Codex implementation and checkpoints, review, bounded QA retries, a pull request, explicit merge, and closeout

Drig workflow diagram

That is the version of the system I can use. I explain what I want to change, inspect the relevant repository, decide on a bounded slice, implement it directly with Codex, and run the checks. Drig remains useful at the edges: clarifying a project request, recording acceptance criteria, preserving checkpoint evidence, and guarding the final QA and delivery boundary.

The direct loop also restores visibility. I can see which question we are answering, what assumption we are making, and why the next command is needed. When a test fails, I can follow the evidence from the failure to the code without first rebuilding an invisible chain of handoffs.

That visibility brought back something I had not expected to miss: the pleasure of coding. There is a particular satisfaction in finding a problem, changing a line, running the check, and seeing the behavior move in the right direction. When too much of that loop is hidden behind automation, the developer loses not only control but also the feeling of making something.

What I keep from the experiment

The experiment did not prove that graphs or harnesses are bad. It gave me a more specific rule: orchestration must justify its own coordination cost.

I keep a separate boundary when the work needs independent permissions, repeatable transitions, a durable acceptance record, or a final decision that should not be made by the implementation process itself. I do not add a new agent because a process diagram has another empty box. For a small, reversible change, one capable model and a visible verification loop are often the better harness.

What stood out when I reread the documentation was that retries have an endpoint. After the first QA fails, the owner chooses whether to fix or stop. A fix allows one more QA run; another failure closes that attempt. Closeout after merge does not repeat review and QA from the beginning. The flow separates stages that need model judgment from stages that record work already completed.

An orderly workflow moving stable inputs through clear stages toward predictable outputs

When inputs, transitions, and outputs are stable, structure can make work easier to understand and repeat.

The questions for the next automation

Before adding another layer, I want to answer five questions:

  1. Is the work repetitive enough to justify orchestration?
  2. Can a person inspect and explain every transition?
  3. Does the automation save more time and context than it consumes?
  4. Can a developer take over without reconstructing hidden state?
  5. Is there a clear point where the system stops instead of adding another recovery step?

I am not giving up on automation. I am giving up on the idea that more automation is always more advanced. For development, the best tool may be the one that keeps the developer close to the code, the evidence, and the decisions.

Dropping the all-in-one version of Drig felt less like abandoning the future and more like leaving a painful detour. The project still taught me where structure helps. More importantly, it let me return to a development loop that I can see, explain, and enjoy.