What are we actually pointing at when we say "harness"?
Is it the product, the orchestration layer, the evaluation scaffold, or just a prompt with retry logic bolted on? And when two engineers on the same team mean different things by the word, is that a vocabulary problem, or a design disagreement they have not yet noticed?
I want to argue the third thing: it is a design disagreement, the disagreement is about determinism, and the field currently has it backwards.
The word genuinely is undefined.
This is not a strawman. A June 2026 paper by Sanderson Macedo [1] is, as far as I can tell, the first work to define "agent harness" head-on rather than in passing, and it opens by documenting exactly how loose the usage has become. The term sometimes denotes a whole product like Claude Code or Codex CLI, sometimes the evaluation scaffold that runs an agent against benchmark tasks [7], and sometimes it gets conflated with an agent framework, an SDK, an IDE plugin, or an orchestrator [1].
The paper also notes that the surrounding literature touches the concept only by facets: agent surveys describe memory, tools and the reasoning loop without consolidating the harness as a unit, agent-computer-interface work covers tools and stops, context engineering covers context and stops [1].
So the data engineer who says "harness means context" and the backend developer who says "harness means long-running calls plus human approval" are both citing real literature. They are each holding one facet.
Macedo's fix is a four-condition membership test [1]. A system is a harness if and only if, at runtime, it has:
T1. An agent loop interleaving reasoning, action and observation.
T2. A tool interface that can alter an external environment, not merely read it.
T3. Active context management, where what enters and leaves the window depends on the task rather than on buffer size.
T4. At least one control mechanism whose effectiveness does not depend on the model choosing to cooperate.
It is a good definition. It is also, I think, drawn from the wrong sample, and its boundary is drawn in the wrong place.
Where the paper and I part ways
Macedo explicitly excludes the orchestrator from the concept [1]. The argument is that a harness is characterized by an adaptive loop where the next step depends on the observation of the previous one, rather than a fixed graph, and that a pipeline which always runs A, then B, then C, without letting observation alter the course, is an orchestrator.
If a harness is an Agent Graph, and a graph is excluded by definition, one of us is wrong.
I think the paper collapses two different properties into the phrase "fixed graph": static topology and fixed sequence.
A pipeline that always runs A, then B, then C is a fixed sequence. Observation changes nothing and the trace is identical on every run. That deserves exclusion. An Agent Graph is not a sequence. The node set and the legal edges are declared before runtime. What happens at runtime is not.
Conditional branching. Routing is a predicate over graph state, and state is written by observations. Two runs of the same graph take different paths and touch different nodes.
Parallel execution. A node can fan out to several successors that run concurrently and join downstream. There is no "next step" at that moment, there is a frontier. The cleanest semantics for this are bulk synchronous parallel super-steps [10], where every vertex in a step reads the same frozen state snapshot and reducers fold results back into state after the barrier in a stable order. Deterministic arrangement, concurrent execution, no race on state.
Cycles. A node can route back to itself or upstream and iterate until a termination predicate holds: a maximum iteration count, a convergence check, a validator passing, or a policy node returning approved.
That last one is where the exclusion breaks down entirely. ReAct [6] is a cycle. Reason, act, observe, route back on the observation. That is a two-node graph with a conditional self-edge and a termination condition. The adaptive loop Macedo makes the defining property of a harness is not the alternative to a graph. It is the smallest interesting graph.
So the definition admits the special case and excludes the general one. And the general case buys something specific. Among the fourteen failure modes Cemri et al. catalogue, missing termination conditions sits in the specification and system design category alongside task misinterpretation, ambiguous role definitions and poor decomposition [2]. A cycle declared as an edge forces you to write the termination predicate down, where it can be reviewed, version-controlled and tested. An implicit loop inside a model's own control flow leaves termination to the model's judgment about whether it is finished, which is close to the last judgment I would want to delegate.
"Deterministic," then, does not describe the trajectory, the ordering, or the timing. It describes the arrangement: which nodes exist, which transitions are legal, what state each node may read and write, and under what condition a loop stops. Non-determinism stays quarantined inside the nodes that hold a model, and there is a named list of those nodes.
An agent that re-derives its own control flow at every step has not eliminated the graph. It has made the graph implicit, unnamed, and different on every run, which means it cannot be tested, replayed, audited or explained.
Why the four conditions are not enough
The deeper issue is the sample. Macedo applies the test to six systems: Claude Code, Codex CLI, Aider, Cline, OpenHands and SWE-agent [1]. All six are coding agents.
Coding agents live in an unusually forgiving domain. They get a free deterministic verifier, the test suite. They have one actor and no counterparties. Their work is reversible through version control. A task starts and ends inside one session. Almost nothing they do requires a second human's signature.
Strip those privileges away and run the same definition against a claims adjudication process, a KYC onboarding flow, or a trade reconciliation, and the four conditions stop being sufficient. They say nothing about work that outlives the session, nothing about approval as a structural element, nothing about what survives between sessions, and nothing about whether the arrangement resembles the business process it is automating.
Those are exactly the node types I would argue constitute a harness.
Agent nodes. Nodes with a model inside. Judgment, extraction, synthesis. The only place stochasticity is licensed. The quality of a harness correlates inversely with how many of these it has.
Long-running tool nodes. Anything whose duration exceeds the request. The industry rediscovered durable execution for this reason, and the sharpest framing I have seen of it is that a workflow definition and a workflow execution are different objects: the definition is the static blueprint of nodes, edges, routes, policies and schemas, while the execution is one instance with a run ID, a current node, persisted state and an event history [8]. That is precisely the static-topology and dynamic-path split, arrived at independently by the workflow engine community two decades before agents existed.
Action policy nodes (HITL). Approval as a node, not a boolean on a tool. A node has state, a timeout, a rejection edge, an identity requirement and an audit record. A flag has none of those. Runtime implementations now treat this as an interrupt that checkpoints full graph state to durable storage and waits indefinitely until a resume command arrives, whether that is minutes or days later [9]. Macedo's T4 asks only for one control mechanism independent of the model [1]. In a regulated domain, control is not one mechanism. It is a topology.
Long-term memory nodes. What survives the session. CoALA [5] is still the cleanest decomposition here, splitting agent state into working memory holding active variables and observations, and long-term episodic, semantic and procedural stores. The important design consequence is that writes to long-term memory should be node transitions, deliberate and inspectable, not a side effect that accumulates.
Session context. What is true for this invocation. The transcript, the working key-value state, the original payload.
Function nodes. Ordinary code. Validation, persistence, transforms, deterministic routing. Most of a good graph is this. Macedo's own paper gestures at it under "deterministic handlers", and concedes something stronger in passing: that the robust response to an agent claiming a success it did not achieve is to verify state deterministically and run sensitive parts as ordinary code, because politely instructing the model not to do so is the weakest available control [1].
If a node can be a function, it must be a function. That is not a performance optimization, it is the containment strategy.
Routing is a function of state, not of the last message
The most common architectural mistake I see is routing on the most recent model output. It works in a demo and it is brittle forever after, because any change to an upstream prompt silently changes downstream routing.
Route on state instead. The state is the accumulated key-value pairs, the session messages, and the invocation. A condition over state is a thing you can unit test, assert on, and replay. A condition over the last generated string is a thing you can only hope about.
The evidence that hoping does not scale is unambiguous. The tau-bench authors introduced pass^k, the probability that an agent succeeds on all k independent trials of the same task, and found that even the best-performing agent with over 60% average task success dropped below 25% at pass^8 [3]. Same task, same tools, same policies, eight runs. Roughly a one-in-four chance it behaves correctly every time. No amount of model improvement fixes a variance profile that comes from unconstrained control flow.
The graph must mirror the domain
This is the part I hold most strongly, and it has the best empirical support.
Cemri et al. [2] built the first grounded taxonomy of multi-agent failures from more than 1,600 annotated traces across 7 frameworks, producing 14 failure modes in three categories: system design issues, inter-agent misalignment, and task verification. Specification and system design accounts for roughly 42% of observed failures, covering task misinterpretation, ambiguous role definitions, poor decomposition, duplicated roles and missing termination conditions. The authors are explicit that these failures stem from system design rather than from model limitations or simple prompt-following, and require more than superficial fixes [2].
Every item on that list is a graph defect. Missing termination conditions are missing terminal nodes. Ambiguous role definitions are nodes without a contract. Poor decomposition is wrong node granularity. Roughly 42% of the failure surface is the arrangement.
The converse result is older. MetaGPT [4] took the position that cascading hallucinations arise from naively chaining LLMs, and encoded standardized operating procedures into the workflow so that agents with domain-specific roles verify intermediate results and reduce errors. Its own account of the mechanism is that every handover must comply with an established standard, and that structured intermediate outputs maintain consistency and minimize ambiguity during collaboration [4].
The SOP is the domain process. Encoding it was the intervention. This is domain-driven design (DDD) with different vocabulary:

Figure 1. This is domain-driven design (DDD) with different vocabulary.
You are not inventing governance. You are transcribing it.
Which yields a falsifiable test. Put the graph in front of a domain expert who does not write code. If they can read it and correct you, and say something like "no, legal reviews before pricing", you have a harness. If they see one box labelled "agent" with arrows to everything, you have a chatbot with tools and an expensive incident waiting.
Back to the questions
What are we pointing at when we say harness?
A graph: agents, long-running tools, action policies, long-term memory, session context and function nodes, connected by edges whose traversal is a function of state.
Is the disagreement terminological? No. The data engineer defending context and the backend developer defending long-running calls and HITL are each defending one node type. The argument dissolves once you see the graph, because both are in it.
And is the field wrong? On one specific point, I think so. The best current definition excludes the orchestrator on the grounds that a fixed graph is not adaptive [1]. But static topology and fixed sequence are different properties, and once a graph can branch, fan out and cycle, the excluded category turns out to contain the included one: the adaptive loop is a two-node graph with a self-edge. A definition that rules out the general case while admitting its own special case is not drawing a boundary, it is drawing a preference.
The model is rented and it changes every quarter. The graph is yours. It is worth defining precisely.
References
[1] S. O. de Macedo, "What makes a harness a harness: necessary and sufficient conditions for an agent harness", arXiv:2606.10106, June 2026.
[2] M. Cemri, M. Z. Pan, S. Yang, L. A. Agrawal, B. Chopra, R. Tiwari, K. Keutzer, A. Parameswaran, D. Klein, K. Ramchandran, M. Zaharia, J. E. Gonzalez and I. Stoica, "Why Do Multi-Agent LLM Systems Fail?", arXiv:2503.13657, 2025.
[3] S. Yao, N. Shinn, P. Razavi and K. Narasimhan, "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", arXiv:2406.12045, 2024.
[4] S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, C. Zhang, J. Wang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu and J. Schmidhuber, "MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework", arXiv:2308.00352, 2023 (ICLR 2024).
[5] T. R. Sumers, S. Yao, K. Narasimhan and T. L. Griffiths, "Cognitive Architectures for Language Agents", arXiv:2309.02427, 2023.
[6] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan and Y. Cao, "ReAct: Synergizing Reasoning and Acting in Language Models", arXiv:2210.03629, 2022.
[7] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press and K. Narasimhan, "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?", arXiv:2310.06770, 2023.
[8] Koshy, "Agent Workflows Are Rediscovering Durable Execution", Beyond Localhost, May 2026.
[9] LangChain, "The Runtime Behind Production Deep Agents", April 2026.
[10] G. Malewicz, M. H. Austern, A. J. C. Bik, J. C. Dehnert, I. Horn, N. Leiser and G. Czajkowski, "Pregel: A System for Large-Scale Graph Processing", SIGMOD 2010.
As of 21 September 2026. References [1] to [7] and [10] are peer-reviewed or preprint literature; [8] and [9] are engineering write-ups, cited as practitioner evidence rather than as findings.
Join the conversation! Your thoughts help the community grow.