AI harness

Definition
The tools, context, rules and memory built around an AI model that turn it from a chat box into a system reliable enough to run in production.

Why it matters

Two teams can use the same model and get very different results. The difference is rarely the model. It is whether one team gave it the right tools, trimmed the context to what matters and added a check before anything reaches a customer. When an AI workflow disappoints, the setup around the model is usually the first place to look.

What it is

A language model on its own only turns text into text. To do useful work it needs a setup around it, and that setup is the "harness". It usually has six parts:

  • Instructions that say what the job is and what counts as done.
  • Tools the model can call, such as a CRM lookup or a send-email action.
  • Context, meaning the records and history it is shown for this task.
  • Memory of what happened in earlier runs.
  • Rules and checks that stop bad actions or catch bad output.
  • A place to put the result, such as a record, a queue or a document.

Claude Code is a good example. The model is Claude, and the product around it, with file access, a terminal, permissions and a rules file, is the "harness".

Common mistakes

Giving a model every tool at once, so it picks the wrong one. Stuffing in every document it might need, which buries the useful part. Skipping the check because the first ten runs looked fine.

How to build one

  • Start with one job and write down what a good result looks like.
  • Add only the tools that job needs.
  • Show the model the relevant record and recent history, not the whole system.
  • Add a check before any output touches a customer or a system of record.
  • Keep a log of every run, so mistakes can be found and fixed.
Worked example

Suppose a twelve-person B2B services firm wants a model to draft first replies to inbound enquiries. The first attempt gives the model the whole CRM and an action that sends email. The drafts are inconsistent, and some quote the wrong price. The team rebuilds the harness around one job. Their code calls the Anthropic API with the enquiry and its matching company record only. The model may draft but not send. A check rejects any draft quoting a price that is not on the record. Each run is written to an Airtable log with its inputs and output, so a wrong draft can be traced. Say one draft in ten is rejected in the first week. The fix came from the context and the check, not from a different model.

Tools in the example

Some links are affiliate links: we may earn a commission at no cost to you. It never decides a ranking. How we work with partners

  1. Article

    System prompt

    The standing instructions in the setup.

  2. Article

    Function calling

    How the model actually uses a tool.

  3. Article

    Guardrails

    The rules that keep its actions safe.

  4. Article

    Context engineering

    Choosing what the model sees.

  5. Article

    Human-in-the-loop

    The review step in the flow.