Harness Engineering
Ask which model a team uses and you learn surprisingly little about whether their agent works. The model is the reasoning — brilliant, but with no hands, no memory of yesterday, and no way to touch the world. Everything that turns that reasoning into reliable action sits in the layer around it: the harness. A common rule of thumb puts the model at roughly a tenth of what makes an agent actually useful, and the harness at the other ninety per cent.
Harness engineering is the craft of building that layer well. The term was coined by Mitchell Hashimoto — founder of HashiCorp and co-creator of Terraform — in February 2026, from a simple habit: every time an agent made a mistake, he engineered a permanent fix into the agent’s environment rather than just re-prompting it. Within weeks OpenAI and Anthropic had published their own takes, and the phrase stuck. This page explains what a harness is, everything it is responsible for, why it out-weighs the model, and how it sits inside the wider craft of prompt engineering.
What harness engineering actually is
Harness engineering is the practice of designing the runtime substrate around an autonomous agent — the tools it can call, the memory it keeps, the sandbox it runs in, the checks on its work, and the limits on what it may do — so that a single agent can run for a long time, on its own, and stay both safe and reliable.
The neat way the field states the relationship is a formula:
The percentages are a rule of thumb, not a measurement — but the point they make is real. A better model raises the ceiling on a single thought. A better harness is what lets that thought survive contact with real tools, real files and real consequences, over and over, without a human watching every move.
Where the harness sits
The whole craft of steering a model — prompt engineering in the broad sense — can be drawn as nested rings. At the core is the prompt: the context, skills and wording for a single model call. Wrap that in a harness and you have one capable agent that can act safely and reliably. That harness — the ring around the core below — is what this page is about.
Keeping the harness in focus settles a common muddle: the harness is not a security wrapper for a whole fleet of agents — it is what makes a single run safe and capable. Running a harnessed agent over and over, unattended, is a wider concern again, and it has its own page next to this one.
You don't need to build this from scratch if you are using purpose-built agentic CLI tools like Claude Code. These environments come with a built-in base harness out of the box—handling local context collection, execution loops, tool definitions, and baseline permission prompts for file edits or bash commands. For individual workflows, you are often configuring and tuning an existing harness rather than engineering one from the ground up.
In one widely-cited experiment, the same underlying model went from producing work that “technically launches but is broken” to something fully functional — not by swapping the model, but by adding a separate evaluator to the harness so the agent could no longer grade its own homework. The lesson repeats everywhere: architectural constraints and a well-built environment out-perform raw model capability for sustained, unattended work. Choosing a smarter model is the easy lever; engineering the harness is the one that actually decides whether the agent is trustworthy.
The core habit: fix the environment, not the prompt
Hashimoto’s original insight is worth stating plainly, because it is the method. When an agent gets something wrong, you have two choices. You can re-prompt it — nudge it in words, this once, and hope it remembers next time. Or you can ask a sharper question: what about the environment let this mistake happen, and how do I make it impossible?
The second question is harness engineering. A missing convention becomes a lint rule the agent cannot pass without following. A dangerous command becomes a permission it does not have. A step it keeps forgetting becomes a checklist baked into the tools. A vague sense of “done” becomes a test that has to go green. Each fix is permanent and mechanical, so the same mistake cannot recur — and, crucially, the fix helps every future run, not just this one. Re-prompting scales with your attention; engineering the harness compounds without it.
What a harness is responsible for
‘Tools and a system prompt’ barely scratches it. A serious harness carries most of what makes autonomous work trustworthy. These are the parts worth designing on purpose — expand each to see what it does and why it earns its place.
-
The instructions that are always present before any user input — who the agent is, what it may and may not do, the conventions it must follow. This is the one part that overlaps most with classic prompting, and it sets the baseline for every run.
-
The functions the agent can actually call: APIs, databases, code runners, shell commands, a browser. A model can only propose an action; the harness is what lets it take one, and defines exactly which actions are on the menu.
-
An isolated workspace where the agent can run code and touch files without affecting anything outside it. This is the security foundation: the blast radius of a bad step is bounded by the sandbox, not by the agent’s good judgement.
-
Three kinds, working together: short-term (what fits in the context window), working (a scratchpad for the current task), and long-term (files, a database or a vector store the agent can search). The harness decides what to load, what to keep, and what to drop before the context overflows.
-
The harness can run tests, inspect real output, and prompt the model — or a separate evaluator agent — to review the work before it is accepted. Keeping the reviewer distinct from the author is what stops “the model says it is done” being mistaken for “it is done.”
-
Rules that block unsafe actions outright, and checkpoints that pause for human approval before anything sensitive or irreversible — a deletion, a payment, a production deploy. Autonomy is granted deliberately, one action type at a time, not assumed by default.
-
A full trace of what the agent saw, decided and did — prompts, tool calls, results, tokens and cost. Without it you cannot debug a failure, attribute a cost, or satisfy anyone who needs to audit the system. You cannot trust what you cannot see.
-
What happens when a step fails: retry, back off, escalate to a human, or stop. And, just as important, knowing when to stop at all — a hard limit on time, cost or attempts so a stuck agent cannot spin forever or burn a fortune chasing a goal it cannot reach.
The vocabulary this needs
Terms worth agreeing on before you design a harness.
How it relates to prompt engineering
Harness engineering is not a rival to prompt engineering — it is prompt engineering, taken in its broad sense (the craft of getting useful behaviour out of an LLM through everything you feed it) and applied at a wider scope. Two layers sit inside it:
- Prompt / context, at the core — the wording, examples and information for a single model call. A harness is stuffed full of prompts: the system prompt, the tool descriptions, the evaluator’s brief.
- Harness, around one agent — this page. Everything that makes a single run safe, capable and reliable.
The boundary to hold in your head: the harness works per run. Everything here is about making one agent’s run trustworthy — the model still does the reasoning, and the harness turns it into safe, repeatable action. Running that harnessed agent many times, on a schedule and without you watching, is the next scope out, and it has its own page beside this one.
Building one: where to start
You do not design a whole harness up front; you grow one, mistake by mistake. A practical order:
- Start with a sandbox. Before anything ambitious, make sure a wrong step cannot reach anything you care about. Isolation first.
- Give it exactly the tools it needs — no more. Every tool is a capability and a risk; a smaller menu is easier to reason about.
- Make ‘done’ mechanical. Add a test, a linter or a schema the agent must satisfy, so success is something the harness can check rather than something the model claims.
- Add a separate evaluator once the work matters. Do not let the agent that did the work be the one that signs it off.
- Instrument everything. Log prompts, tool calls, results and cost from day one — you cannot improve a harness you cannot see into.
- Then, every time it errs, fix the environment. Turn each mistake into a permanent, mechanical constraint. That habit, repeated, is the harness.
The model is the tenth of the agent you cannot control; the harness is the nine-tenths you can. Reliable agents are not the ones running the cleverest model — they are the ones whose environment makes the right thing easy and the wrong thing impossible. When an agent disappoints you, resist the urge to re-prompt. Ask what the environment allowed, and fix that.
Sources and further reading
This page is our own synthesis of an idea that took shape publicly in early 2026, following Mitchell Hashimoto’s original framing and the engineering write-ups that followed.
- Milvus, What Is Harness Engineering for AI Agents? — milvus.io/blog
- Databricks, What is an AI Agent Harness? — databricks.com/blog/ai-harness
- MindStudio, What Is the Agent Harness? — mindstudio.ai/blog
- Firecrawl, What Is an Agent Harness? — firecrawl.dev/blog
Read more
The three phases
CRAFT
Craft (write) the prompt with the following elements: Context, Register, Acting Role, Format, and Task.
ING
Validate the prompt and ensure it maintains an interactive approach. Keep in mind the importance of non-disclosure and staying goal-driven throughout the process.
AI
Continuously assess and refine the output based on the prompts output to improve the overall quality.