One of the common mistakes when building an AI coding workflow is asking one model to interpret the requirements, implement the code, run the tests, review the changes, and decide whether the work is ready to merge. That is convenient for a small task. As the work grows, context gets longer and it becomes harder to tell why something failed or who was responsible for the decision.

My answer is not to add as many agents as possible. It is to give each role a boundary that a person can inspect. The useful distinction is not “how many models are running?” but “what may this role decide, change, and approve?”

The role map has to be inspectable

The current public Drig lifecycle makes the separation concrete. Its model-backed stages and deterministic stages are named independently:

Stage Route documented by Drig Responsibility
ideaInterview gpt-5.6-terra / high Clarify only blocking scope gaps
internalPlan gpt-5.6-sol / high Turn the approved context into a bounded plan
Development and checkpoint gpt-5.6-sol / high Implement a focused slice and leave evidence
codeReview gpt-5.6-sol / high Review each immutable checkpoint
QA revision 1 and 2 gpt-5.6-luna / max Check unresolved requirements and evidence
Routing, validation, publication, closeout No model route Enforce deterministic contracts

The exact model names can change. The durable part is that the route is explicit. Drig exposes model-route --stage <stage> so a stage does not silently receive a default model when its policy is missing. Unknown stages fail closed.

A small task is divided into planning, implementation, review, and QA lanes, each with a visible boundary and evidence card

The role boundary is more valuable than the number of agents.

Permission is part of the role

A role is incomplete if it only names a model. It also needs a permitted action and a stopping point.

An implementation role may edit the assigned slice, run the checks that belong to it, and leave a local checkpoint. It should not quietly turn its own code into final approval. A review role should be able to inspect the change and record a finding without modifying the product as a side effect. A QA role should receive the acceptance criteria and bounded evidence it needs, not an unlimited copy of every previous conversation.

That separation is why the Drig lifecycle keeps one durable slice worktree per checkpoint, runs a native review per checkpoint, and gives requirements QA an evidence manifest. The final publication and merge decisions remain separate from the model that wrote the code. This is slower than letting one agent announce its own success, but it makes a wrong result easier to locate.

A small task should remain small

Nested roles are not a badge of maturity. A one-file, reversible change may need one capable Codex session and a direct test run. Splitting that task into planning, implementation, review, and QA agents adds handoffs without creating a useful boundary.

I add a second role when at least one of these conditions is real:

  • the work needs a separate permission set;
  • the acceptance criteria should be checked by a process that did not write the change;
  • the operation repeats often enough to justify a durable contract;
  • a failure must be resumed from a checkpoint instead of reconstructed from chat;
  • the owner needs a separate approval point before publication or merge.

If none of those conditions exists, the extra model is probably ceremony.

What I kept after simplifying the harness

I removed the heavy wrappers and kept the part that helped: role-specific context, explicit model routing, focused workspaces, and evidence that another role can read. The current Drig plugin is therefore not a collection of autonomous agents competing to finish the same task. It is a workflow where ordinary Codex does product development and the surrounding stages make planning, review, QA, and delivery visible.

This also changes how I evaluate a model. I do not ask only whether it can write code. I ask whether it can stay inside a slice, leave a useful checkpoint, respond to a validator, and stop when its role ends. A strong implementation model can still be the wrong QA model if it encourages the wrong kind of confidence.

What this structure does not solve

Role separation does not make a workflow automatically correct. More stages still cost time and context. A badly written acceptance criterion can be checked very carefully and still describe the wrong outcome. A model route can be explicit and still be poorly chosen.

The structure only gives those failures a place to appear. That is enough to make the next change more deliberate. The goal is not a larger team of agents. It is a smaller set of responsibilities that a person can explain from the request to the final evidence.

That is the principle I mean by Nested Codex. Use one model when one model is enough. Add a role when it creates an independent boundary, and make its model, permissions, evidence, and stopping point visible from the beginning.