Model names and benchmark scores are the easiest way to talk about AI agents. They produce a clean ranking: one model is stronger, another is cheaper, and a third is better at coding. But a model does not complete a real task by itself. It works through an execution layer that supplies context, tools, permissions, state, retries, and stopping rules.
That layer is often called the agent harness. It is easy to overlook because the final answer still appears to come from the model. In practice, the harness decides what the model can observe, what it can change, and whether it can recognize that an apparently successful run has produced the wrong artifact.
Harness-Bench, a diagnostic benchmark introduced in a May 2026 paper, makes this layer explicit. The researchers evaluated 106 sandboxed offline tasks and analyzed 5,194 execution trajectories across model-harness pairings. Each run recorded not only completion, but also final artifacts, execution traces, usage statistics, and validator outputs. They report meaningful variation in completion, process quality, efficiency, and failure behavior.

The execution path around a model can change both the result and the way failure appears.
The model–harness pair is the unit of evaluation
The paper does not show that one harness is always best. Its more useful claim is that agent capability should be reported at the model-harness configuration level rather than attributed to the base model alone. The authors also identify execution-alignment failures: the model’s reasoning can become detached from tool feedback, workspace state, evidence, or the contract used to verify the output.
That claim gives a concrete reader question: what must be held constant before two agent results can be compared? At minimum, the task, initial input, model version, context policy, available tools, permissions, time limit, retry budget, and validator. Change any of these, and the observed capability may change with it.
A fair comparison needs a harness manifest
An evaluation report should expose the conditions that can move the outcome:
| Field | Why it matters |
|---|---|
| Model and version | Model updates can change behavior even when the task is unchanged. |
| Task and initial input | A hidden setup step can be as important as the visible prompt. |
| Context and memory | Retrieved instructions may help one run and distract another. |
| Tools and permissions | A read-only agent is solving a different problem from an agent allowed to edit and execute. |
| Time, retry, and token budget | Recovery opportunities can turn a near miss into a pass. |
| Validator and output contract | A permissive checker can inflate success; a brittle one can hide useful partial work. |
This is also why traces matter. A final answer cannot distinguish a correct plan, a successful tool call, a failed call followed by recovery, and an answer that merely passed a weak validator. The useful artifact is the result together with the trace, tool outcomes, usage, validator decision, and failure reason.
In a repository workflow, the manifest does not need to be a large report. To compare two runs of the same bug-fix task, hold the model version and initial repository state constant, then record whether each run was allowed only to read and test or also to write a patch. That separates a model effect from an authority effect. The same record should say whether the network was available, how many failed tests could be retried, and which check accepted the final file.
This record matters even more when a run fails. A success rate can invite the conclusion that a stronger model is needed when the actual cause was a missing tool, excessive authority, a short budget, or a validator error. Failure labels make it possible to change one variable at a time in the next run.

A comparison remains meaningful only when the tool, context, authority, budget, and validation conditions travel with the model name.
Failure should be split before it is scored
Agent failures are not interchangeable. A model may misunderstand the task, lose a required fact from context, lack permission for the necessary action, call a tool with the wrong arguments, retry unsafely, or produce a good artifact that the validator cannot recognize. Counting all of these as “model failure” points the fix in the wrong direction.
The same separation applies to orchestration systems. Drig’s public workflow distinguishes planning, ordinary implementation, code review, QA, and deterministic delivery stages. That does not prove that every workflow needs more agents. It shows a useful evaluation boundary: identify which stage owns the decision, what evidence it must leave, and which transitions are allowed to fail closed.
What to compare in practice
Before comparing two models, freeze the harness or publish the difference. If the goal is to compare models, keep tools, context, permissions, budgets, and validators constant. If the goal is to compare harnesses, keep the model and task set constant and report the extra retries, retrieved context, tool surface, and recovery behavior.
For an individual project, a small harness manifest is usually more valuable than another headline score. Record the input, model version, available tools, permission boundary, memory policy, retry limit, final artifact, and validator output. When a run fails, label the failure before changing the prompt. That turns an anecdote into a diagnosis.
There is no local controlled benchmark behind this article, so the recommendations here are a reading of the cited benchmark and public workflow documentation, not a claim about a new measured win. The next useful experiment is a paired run in which one variable changes at a time: first the model, then the tool boundary, then memory, then recovery policy. Only then can a score say what caused the difference.




