In the first part of this series, I wrote about leaving an OpenCode multi-agent setup behind and moving to Codex. I thought I would put harnesses aside and just get on with development. GPT worked well for me, and my usage allowance felt generous.

Then the market kept moving. New tools and models appeared so quickly that I started to feel I was falling behind whenever I ignored one. Looking back, it was mostly FOMO. “Graph engineering” sounded like something beyond a harness. When I looked closer, it often turned out to be a prettier way to wrap work I was already doing in Codex.

The Codex app was different. Its updates kept making the interface easier to use, and the project view fit the way I worked across several codebases. While I was looking around, Codex was quietly getting better.

I wanted a personal assistant

I also had work I wanted to hand off beyond coding. That was when Hermes Agent caught my attention. The split seemed simple: Codex would handle development, and Hermes would act as a personal assistant.

Hermes would organize conversations and requests, route each command to a different model, and send development work to Codex through codex exec. It would then collect the result. If it worked, it might feel like a small team. Self-learning was one of the reasons I chose Hermes, and the Slack and Discord integrations made the idea even more attractive.

The trouble started as soon as the two tools had to work together. Hermes could not really see what was happening inside Codex. The middle of the process became a black box. I kept adding explanations and summaries so both tools could hold on to the same context. Before long, I was spending more time making the tools understand each other than getting work done.

The same context was being read and summarized twice, so the handoff consumed tokens on both sides. OpenRouter added another bill on top of the Codex subscription I already had. Later, when I could follow and continue Codex work through the ChatGPT mobile app, Slack and Discord mattered less to me. I went back to using Codex on its own.

Task cards travel between two computers on a long conveyor, with a few tokens falling below

Splitting the tools also meant managing the explanations and results moving between them.

Then I built Drig

The idea of an agent team did not go away. I took the parts I liked from Hermes and started building Drig myself. The plan was to connect planning, interviews, development, QA, review, and release in one graph.

I even intended to open-source it, so I spent more than a month working on it. At some point, though, I was spending more time fixing Drig than doing the development work it was supposed to support.

Review and QA were where it fell apart. I tried to process only the changed parts, but the review list and fix list kept showing the same work again. I would clear an item, only to see it return in the next step. I should have removed the old route as I added the new one. Instead, the old steps stayed and new branches kept accumulating.

The graph grew, and so did the verification path. A ten-minute change could take an hour to verify. At the time, I thought this was an AI problem: solve one issue while leaving the old ones in place.

Looking back, I was doing the same thing. I complained that the AI kept old routes, then kept those routes myself and added another control step on top. When Codex Sol arrived, Drig felt heavy and indecisive. I removed it from my projects.

This is a recollection of how the workflow felt, not a controlled comparison of time or cost, and it is not a verdict on any particular tool.

What I will check first next time

This experience changed how I choose tools. I now care less about how many roles a workflow can split into and more about how much work I will have to manage after the split.

  • Is the handoff necessary? If the next tool has to summarize the previous tool’s result again, I first look for a way to remove that handoff.
  • Can I see the current state? I need to see not only what was done, but also why the work is blocked and what should happen next.
  • Does a new step remove an old one? When I add a fix, I should be able to delete a route that is no longer needed instead of keeping every branch.
  • Am I counting verification as a cost? Saving development time is not enough if review, QA, and fix-list management create a larger bill afterward.

After years of working on development teams, I still want agents to work together like an ideal team. For now, Codex stays at the center. If I build another graph, I want to count what it actually removes—and what new verification work it creates—before I draw the impressive structure.