The AI Model Isn’t the Moat Anymore. The Harness Is. What Is AI Harness Engineering?

The AI Model Isn’t the Moat Anymore. The Harness Is.

The next advantage in AI will come from how reliably an agent can remember, act, recover, and prove what it did.

A customer asks an AI agent to review a sales account. The agent checks the CRM, reads recent support tickets, finds a contract in a document store, and drafts a recommendation. Halfway through, one tool times out. Another returns an outdated record. The contract contains instructions the agent should treat as document text, not as orders. Before the recommendation reaches the customer, someone needs to know which sources it used and whether it made any changes.

The model matters at every step. But a model alone cannot decide which systems it may access, preserve the task through a failure, or provide a trustworthy record of its actions. The software that manages those responsibilities is the agent harness.

For years, the AI race was easy to describe: whose model writes, reasons, or codes best? That question still matters. Yet as companies move from impressive demonstrations to agents that do real work, another question becomes just as consequential:

What surrounds the model when it has to operate in the real world?

From answers to operations

A chatbot can give a useful answer in one exchange. An agent may need to work for an hour, cross several systems, ask for permission, recover from errors, and return with a result someone can check.

The harness runs that process. It gives the model tools, passes tool results back, manages the working context, enforces boundaries, and decides what happens when a step fails. Anthropic describes a harness as the system that enables a model to act as an agent; evaluating an agent means evaluating the model and harness together.

Think of a talented employee on their first day. Intelligence helps, but so do access to the right files, clear authority, a reliable handoff process, and a way to find out what went wrong. An agent needs the software equivalents.

Those equivalents are where much of the practical advantage now lives:

  • Memory keeps the goal, relevant facts, and completed work available across a long task. Good memory also knows what to discard; repeatedly sending every previous message back to a model can raise cost and bury the facts that matter.
  • Tools let the agent read, calculate, search, create, and update. Their design determines whether it can do useful work without receiving broad, unnecessary access.
  • Retries and recovery let work continue after a timeout or failed environment. A retry must also recognize whether an action already succeeded, so an agent does not send the same email or place the same order twice.
  • Permissions define the line between reading information, preparing an action, and carrying it out. The line should become stricter as the potential consequence rises.
  • Routing sends different tasks to appropriate models and tools. A routine lookup may need less compute than a difficult judgment.
  • Observability records the path to the result: the sources consulted, tools called, approvals received, errors encountered, time spent, and cost incurred.

A strong model in a poorly designed harness can be expensive and unreliable. A well-designed harness can make the same model more useful.

TrueForge versus Claude Managed Agents

The emerging competition is visible in two approaches.

Claude Managed Agents is Anthropic’s hosted service for long-running agent work. Anthropic separates the agent’s session record, the harness that runs the model and routes tool calls, and the sandbox where code and file operations happen. If a sandbox or harness fails, the system can recover from the durable session record. That architecture also creates clearer boundaries around execution.

TrueForge, released by TrueFoundry as an open-source, model-neutral harness, takes a different product approach. Teams can run it on their own infrastructure, bring their own models and tools, and use its agent loop, session persistence, context management, sandboxes, and approval steps. TrueFoundry also offers its gateway for centralized controls such as budgets and traces.

The choice is not simply “managed versus open source.” It is a decision about who controls the runtime, where the agent runs, how easily models can be changed, and who is responsible for operating the surrounding infrastructure.

TrueFoundry published a comparison using 14 enterprise tasks that required agents to work across CRM, project-tracking, and document tools. In its three-trial averages, TrueForge and Claude Managed Agents each solved 10.7 of 14 tasks using the same Opus 4.8 model. TrueForge’s reported average cost was $8.60 per run, versus $11.80 for Claude Managed Agents. When TrueForge used GLM-5.2, TrueFoundry reported 11.7 of 14 tasks solved at $3.00 per run.

That is an interesting result, particularly because the same-model comparison isolates more of the harness’s effect. It is also TrueFoundry’s own benchmark, on a small set of tasks graded by an LLM judge. It does not establish that one system will be cheaper or more accurate for every company. The sensible next step for a buyer is to run both approaches against their own workflows, including failures, approvals, and the full operating cost.

Still, the comparison makes one point hard to ignore: the number of tool calls and the amount of context carried through each step can materially change the cost of an agent run. TrueFoundry attributes much of its result to a leaner loop, fewer tool calls, and compacting history instead of repeatedly replaying it.

The hidden cost of a bad loop

Suppose an agent needs to inspect 20 customer records. On each turn, it receives its instructions, tool definitions, and an ever-growing account of what it has already seen. The final answer might be only three paragraphs, while the agent has processed millions of tokens getting there.

Now imagine an improved harness that retrieves only the relevant records, summarizes completed steps, and sends a smaller model the routine parts of the job. The customer may receive an equally good answer, with less delay and lower cost.

The opposite can happen too. A harness that compresses too aggressively may erase the one exception that changes the recommendation. Efficiency is valuable only if the result remains correct and the evidence remains available.

That is why cost per successful, reviewable task is more useful than cost per model call. A cheap model invocation does not help if the agent makes an unapproved change or requires a person to redo the work.

The moat is a system that earns trust

“Moat” can sound like a claim that models are interchangeable. They are not. Better models can solve problems that weaker ones cannot, and the right model can simplify the harness it needs. Anthropic itself notes that harness assumptions can become obsolete as models improve.

The more durable advantage is a team’s ability to improve the whole system around those changing models. Does the agent have current information? Can it tell a tool instruction from untrusted text in a document? Does it know when to ask a person? Can it resume after failure? Can the company inspect why it reached a conclusion?

For an AI agency, these questions turn into a better client conversation. Instead of selling “an AI agent that handles support,” ask what the agent may read, what it may change, what requires approval, how success will be measured, and what the team needs to see when something goes wrong. Build one narrow workflow and test it against real exceptions before expanding its authority.

A useful pilot might be an agent that reviews incoming leads, checks the CRM for duplicates, drafts a tailored response, and presents it for approval. The measurable outcome is not that the agent produced text. It is that the right lead received an accurate response faster, with fewer manual steps and a clear record of what happened.

The next generation of AI businesses will still compete on access to powerful models. They will also compete on something customers can experience every day: whether those models are connected to the right knowledge, given appropriate authority, kept on task, and held accountable for the result.

The model supplies intelligence. The harness determines whether that intelligence becomes dependable work.

What Is AI Harness Engineering? The Missing Layer Between a Prompt and a Working Agent

Ask an AI coding agent to build a website, and it may produce an impressive homepage in minutes. Then click around. A signup button goes nowhere. The contact form looks finished but sends no email. A feature described in the original request has vanished entirely.

The model may know how to build each piece. The harder problem is keeping a long project organized, carrying accurate progress between sessions, testing the result, and recognizing when the job is actually done.

That is the problem harness engineering tries to solve.

Prompt, context, and harness: three different jobs

The terms overlap, which is why they are easy to confuse.

Prompt engineering shapes the instructions given to a model. A prompt might tell a coding agent to use a particular framework, follow a design style, and explain its changes clearly.

Context engineering determines what information the model sees while it works. Instead of stuffing an entire repository or document archive into a prompt, an agent can retrieve the relevant files, search a knowledge base, and bring in tool results as needed.

Harness engineering organizes the work around the model. It controls the sequence of tasks, the tools and environment available, what gets recorded, when a fresh session starts, how progress is checked, and what happens after a failure.

A simple way to remember the distinction:

The prompt says what to do. Context supplies what the model needs to know. The harness helps it keep working until the result has been checked.

A good harness still uses good prompts and carefully selected context. It puts them inside a repeatable process.

Why long tasks expose the weakness

A short AI task may fit comfortably in one conversation. A large project might require hundreds of decisions over many hours. As the conversation grows, the agent must manage what to retain.

Summarizing earlier work can help, but a summary is an imperfect substitute for the work itself. It may say “contact form complete” when the form was only styled. It may omit a requirement that seemed minor at the time. Once that summary becomes the next session’s starting point, the mistake can travel forward as if it were a fact.

This is why an agent can appear busy for hours yet deliver a partially finished project. It needs a reliable way to distinguish attempted, implemented, and verified.

Anthropic described this challenge in its work on long-running agents. Its approach used an initial setup phase followed by coding sessions that made incremental progress and left clear artifacts for the next session. Those artifacts helped the next agent instance understand the project without relying entirely on a compressed conversation.

The power of a disciplined loop

One practical harness pattern is deliberately simple:

  1. Write down the full requirements as small, checkable tasks.
  2. Select one task.
  3. Give the agent the instructions and context relevant to that task.
  4. Implement it.
  5. Test it against an observable completion condition.
  6. Record what changed and what remains.
  7. Start the next task with a clean working context and the saved project state.

The loop continues until the requirements are complete and the finished system has been tested as a whole.

Some implementations of the Ralph coding-loop pattern use a structured requirements file and work through features one at a time. Anthropic has also explored incremental coding with context resets and handoff artifacts. The details vary, but the useful idea is the same: progress lives in the project’s files, tests, and history, rather than only in the agent’s fading conversational memory.

A loop alone is not enough, however. An agent can repeatedly make the wrong change or repeatedly declare success. The harness needs a meaningful test of completion.

For a website, “the agent wrote a contact form” is a weak test. “A visitor can submit the form, the required fields are checked, and the message reaches the intended destination” is a useful one.

What a harness actually contains

A production harness may include a task list, a tool-calling loop, a persistent session record, access controls, a sandbox for code, error recovery, approval points, and logs that show what happened.

These parts work together. A tool lets an agent change a file; a sandbox limits where it can run code. A progress record says which feature was completed; a test checks whether that claim is true. A permission rule allows the agent to draft a customer email while requiring approval before it sends one.

The best design depends on the task. A research assistant and a coding agent need different tools and different completion checks. Adding more agents or more elaborate planning does not automatically make the system better. Anthropic has noted that harness techniques can become unnecessary as models improve, so teams should keep testing which parts actually help.

The business opportunity

For companies, harness engineering changes the question from “Which AI model should we buy?” to “What work can we make reliably repeatable?”

Consider a lead-response agent. A one-shot prompt might generate a polished email. A useful harness would also check the CRM for duplicates, retrieve the right service information, draft a response using current pricing, route a sensitive lead to a person, and record the outcome. Its success would be measured by correct, timely responses—not by how convincing one draft sounds.

The same principle applies to an AI agent that maintains a website, researches a market, reviews support tickets, or updates a knowledge base. Each needs a clear goal, controlled access, durable progress, and a way to verify the result.

The exciting shift is that capable models can now take on much larger jobs. The engineering challenge is to give them a working environment that supports those jobs from the first step to the last.

A prompt can start the work. A harness is what helps the agent finish it.

Alphire
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.