The GenAI Field Guide

Harness engineering

Design the runtime around the model so work can be performed and checked.

A language model on its own only maps text to text. Everything that turns it into a system that does work lives around it: the loop that calls it, the tools it can use, the sandbox where its code runs, the permissions that limit it, the state that survives a crash, the retries and budgets that keep runs healthy, the checks that decide whether an action may proceed, and the logs that explain what happened. That surrounding software is the harness, and for agentic products it often matters as much as the choice of model. The same model can fail most tasks in a thin harness and succeed at most of them in a good one.

The mental model for this chapter is the model proposes, the harness disposes. Treat the model as a capable but untrusted and nondeterministic component that suggests the next action. Deterministic code then decides whether that action is allowed, runs it in a contained place, records it, verifies the result and decides whether to continue. Prompts steer behaviour; the harness guarantees it. Anything that must hold every time, such as a spending limit, a forbidden command or an approval rule, belongs in harness code rather than in instructions the model may not follow.

The questions build from parts to whole. The Basic questions define the harness and separate it from context engineering, then cover the core parts: the agent runtime, sandboxes, retry policies, checkpointing and state persistence. The Advanced questions cover the controls that make a harness safe to run unattended: execution budgets, permission design, graceful degradation, audit trails, verification gates and lifecycle hooks, then how to keep side effects exactly-once across retries and resumes, how to test the harness itself, and what reliability means for the whole system.

What is harness engineering?

Harness engineering is designing the software around a model that lets it do real work safely: the loop that calls it, the tools it can use, where code runs, what it may touch, how state survives, and how results are checked.

A language model only turns input text into output text. It cannot run a test, read a file, remember yesterday or stop itself from spending money. Everything that turns a model into a working system lives outside it, in what practitioners call the harness: the control code that sends requests to the model, executes the tool calls it proposes, feeds results back, enforces limits and decides when the task is finished.

The harness has a handful of recurring parts. A loop (call the model, run tools, repeat). A tool layer with schemas, permissions and credentials. An execution environment, usually a sandbox, where code and commands run. State that persists across steps and restarts. Policies for retries, budgets and approvals. Checks that verify output before it has effects. And telemetry: traces and an audit trail of everything the agent did.

Why it matters: the same model can look brilliant or useless depending on its harness. A coding agent that can run the test suite, read the error and try again solves far more tasks than the same model asked to write a patch blind. Teams often find that adding a verification step, a better tool, or a hard stop condition improves outcomes more than switching to a stronger model. The harness is also where safety lives, because a prompt can ask a model to behave, but only code can stop it from deleting a production table.

The useful mental model is the model proposes, the harness disposes. Treat the model as a capable but untrusted and nondeterministic component. It suggests the next action; deterministic code decides whether that action is allowed, executes it in a controlled place, records it and checks what happened. Harness engineering is the discipline of designing that deterministic layer well.

How is harness engineering different from context engineering?

Context engineering decides what information the model sees on each call. Harness engineering is wider: it also controls what the model can do, where actions run, how execution proceeds, and how results are checked. Context assembly is one job the harness performs.

Context engineering is the work of choosing and arranging the tokens in each model call: instructions, retrieved documents, tool results, memory, conversation history. Its question is, given a limited window, what does the model need to see right now to choose well? Its failure modes are missing information, irrelevant noise and stale facts.

Harness engineering covers the whole runtime around the model. It includes context assembly, but also the loop, tool execution, sandboxing, permissions, retries, budgets, persistence, verification and logging. Its question is broader: how does a sequence of model decisions turn into safe, finished, checkable work? A test runner, a retry policy, a 30-minute timeout and an approval step for payments all belong to the harness, and none of them is about what the model reads.

The two meet at the boundary of each step. The harness runs a tool, gets 40,000 lines of output, and must decide what goes back into context: the whole thing, the last 200 lines, or a summary plus a file path. That decision is context engineering performed by harness code. Likewise, compaction (summarising old turns when the window fills) is a context technique the harness triggers at the right moment.

A practical way to tell them apart: context engineering changes what the model is likely to decide; harness engineering changes what is possible and what is enforced. If you want the model to prefer the search tool, that is context. If you want it to be impossible to call the delete tool without approval, that is harness. Teams that try to solve harness problems with context ("never delete files" in the prompt) get rules that hold most of the time, which is not good enough for actions that cannot be undone.

What is an agent runtime?

An agent runtime is the program that runs the agent loop: it calls the model, parses tool requests, executes them under policy, feeds results back, tracks state and limits, and decides when the run stops.

The model never executes anything. When it wants to read a file, it emits a structured tool call (a name plus arguments). The runtime receives that, checks it is allowed, runs the real function, captures the result or the error, appends it to the conversation, and calls the model again. This repeats until the model returns a final answer, a limit is hit, or a check decides the work is done.

Around that core loop, a runtime takes on a standard set of jobs. It owns the message history and decides how it is compacted. It holds the tool registry, including schemas the model sees and implementations the model does not. It enforces budgets for steps, tokens and time. It applies retry policies to transient failures. It persists state so a run can resume, and it emits traces for every model call and tool call. Some runtimes also run independent tool calls concurrently and stream partial output to a user interface.

Runtimes come in three broad forms. You can write your own loop, which is a few dozen lines to start and gives full control. You can use an agent framework or a provider's agent SDK, which ships the loop, tool plumbing and tracing. Or you can run agents on a durable execution platform that persists every step so a crash resumes mid-run. Which one fits depends on run length and stakes more than on features: a 20-second assistant turn needs little, a 6-hour job that touches billing needs durability.

Whichever you choose, keep the runtime's decisions deterministic and visible. The model's output is an input to the runtime, never an instruction it must obey. Unknown tools, malformed arguments and requests beyond permission should become error messages returned to the model, not exceptions that crash the run or, worse, actions that execute anyway.

What is a sandbox?

A sandbox is an isolated environment where agent actions, especially generated code and shell commands, run with tightly limited access to files, network, credentials and compute, so mistakes or malicious instructions cannot reach real systems.

Agents that run code are executing text written by a model that may have been steered by anything it read: a web page, an email, a README in a cloned repository. You cannot review every command in advance, so you limit what any command could possibly do. That is the job of a sandbox: a boundary that contains the blast radius of a bad action.

A good sandbox constrains several things at once. Filesystem: only a working directory, often a fresh copy, never the host's home directory. Network: no egress by default, or an allowlist of the package registry and the APIs the task needs, which blocks most data-exfiltration paths. Credentials: none inside the sandbox; the harness makes authenticated calls on the agent's behalf. Resources: CPU, memory, disk and wall-clock limits so a fork bomb or infinite loop dies quietly. Lifetime: ephemeral, destroyed after the task, so nothing persists between users.

Isolation strength varies. A restricted OS user or process sandbox is cheap but shares a lot with the host. A container isolates filesystem and processes but shares the host kernel, so a kernel exploit can escape. A user-space kernel or a lightweight virtual machine (a microVM, which typically boots in well under a second) adds a stronger boundary at a little more cost. A separate cloud VM or a hosted code-execution service puts the risk on someone else's hardware. Choose by who writes the code and what the host can reach.

A sandbox is not a full security model. It limits what code can do inside the box, but tools that the harness exposes (send email, call the billing API) act outside it and need their own permission checks. Treat the sandbox as one layer alongside least-privilege tools and approvals.

What is a retry policy?

A retry policy is the explicit rule for which failures get retried, how many times, with what delay, and at which layer, so transient errors are absorbed without duplicating side effects or multiplying load during an outage.

LLM systems fail transiently all the time: rate limits (HTTP 429), overloaded or unavailable providers (5xx), network timeouts, and dropped streaming connections. Most of these succeed on a second attempt a moment later. A retry policy turns that into a rule rather than an accident: which errors count as retryable, the maximum attempts, the delay between them, and a total deadline.

The standard delay is exponential backoff with jitter: wait roughly 1, 2, 4, 8 seconds, each multiplied by a random factor so a thousand clients do not retry in lockstep and hammer the provider at the same instant. If the server sends a Retry-After header, honour it. Never retry errors that will fail the same way again: 400 bad request, 401 or 403 auth failures, a prompt that exceeds the context window, or a content-policy refusal. Retrying those only burns time and quota.

Agent harnesses have several distinct retry layers and each needs its own rule. Transport retries repeat the same request after a network or provider error. Repair retries happen when the model's output fails validation: you send the validation error back and ask for a corrected version, usually once or twice. Task retries rerun a whole step or run after a failure. Stacking them carelessly multiplies: 3 transport attempts times 3 repair attempts times 2 task attempts is 18 model calls for one step.

The dangerous case is side effects. Retrying a read is harmless. Retrying "send email" or "charge card" after a timeout can do it twice, because a timeout does not tell you whether the first attempt succeeded. Write tools need an idempotency key, a unique id per intended action that the receiving system uses to deduplicate, before they can be retried safely.

What is checkpointing?

Checkpointing is saving a run's progress at defined points, including completed steps, decisions and results, so that after a crash, deploy or pause the run resumes from the last checkpoint instead of starting over or repeating side effects.

Long agent runs die for boring reasons: a process restart during deploy, a provider outage, an out-of-memory kill, a laptop closing. Without checkpoints, a 2-hour research job that fails at minute 110 restarts from zero, paying again for every model call and possibly repeating actions it already took. A checkpoint is a durable snapshot written at a known point, enough to rebuild the run's state and continue.

What goes in a checkpoint: the task definition and configuration (including prompt and model version), the message history or a compacted form of it, the step counter and budget consumed, results of completed tool calls, and a record of side effects already performed. The last item matters most. On resume, the harness must know that the invoice email was already sent so it does not send it again.

There is a subtlety specific to LLMs. Model calls are not deterministic, so you cannot rebuild state by re-running from the start and expecting the same decisions. Store the model's outputs, not just its inputs, and on resume replay recorded results instead of calling the model again. This is the same idea durable execution engines use for ordinary code: record each step's result, and on recovery skip any step whose result is already in the log.

Checkpoint granularity is a trade-off. After every step gives the finest recovery but more writes and storage. After each batch or phase is cheaper and usually enough. A good rule is to checkpoint before and after every side-effecting action, and at phase boundaries for read-only work.

What is state persistence?

State persistence means storing a task's working state, such as status, history, intermediate results, approvals and memory, outside the process and outside the model, so work survives restarts, can move between workers, and can be inspected.

A model is stateless: each call knows only what is in that request. A process is temporary: it can be restarted, scaled down or moved. Any state that must outlive a single call or a single process needs to live somewhere durable, usually a database, object store or workflow engine. That is state persistence.

Agent systems carry several kinds of state with different lifetimes. Run state: status (queued, running, waiting for approval, done, failed), step counter, budget consumed. Conversation state: messages and tool results, often large. Artifacts: files, drafts, generated code. Pending interactions: an approval request waiting for a human, which might take days. Long-term memory: facts about a user or project carried across tasks. Each has its own store, retention and access rules.

Persistence is what makes several harness features possible. Resuming after a crash needs it. So does pausing for a human: an agent that needs approval must save everything, release the worker, and pick up where it left off when the approver clicks. Horizontal scaling needs it, because the next step may run on a different machine. And debugging needs it, because a support engineer must be able to open a run and see exactly what state it was in.

Two rules keep persisted state trustworthy. First, the store is the source of truth, not the model's memory: if the model says the order was cancelled, the harness checks the order system. Second, state changes should be explicit transitions, written by harness code, not free text the model edits. A run moves from waiting_approval to approved because a recorded human decision says so, never because the model wrote that it was approved.

How do you set execution budgets?

Set hard limits on several dimensions at once, such as steps, tool calls, tokens, spend and wall-clock time, per run, per user and per tenant, derive them from measured runs, and define what the harness does when a limit is reached.

Agents decide for themselves how many steps to take, so without limits a confused agent can loop for hours, a broad question can fan out into hundreds of tool calls, and a single bad prompt can cost more than a month of normal traffic. An execution budget is a set of ceilings the harness enforces regardless of what the model wants.

Use several dimensions, because each catches a different failure. Step or tool-call count catches loops. Token count catches context bloat, where each step is cheap but the growing history makes later calls expensive (cost per call grows roughly with history length, so total cost of a run can grow quadratically with steps). Spend catches expensive models or large outputs. Wall-clock time catches slow tools and hung calls. Add per-tool limits for risky or costly tools: at most 3 emails, at most 50 search calls. Then layer scopes: per run, per user per day, per tenant per month.

Derive numbers from data, not intuition. Run the agent over a representative eval set, look at the distribution of steps, tokens and time for successful runs, and set the limit around the 95th or 99th percentile with a margin. A limit at the median kills half the legitimate work; a limit at ten times the maximum catches nothing until it is too late. Revisit after every prompt or model change, because both shift the distribution.

Decide what happens at the limit. A soft limit (say 80%) can inject a message telling the model to wrap up and report what it has. A hard limit stops the run, saves state, and returns a clear status such as budget_exhausted with partial results, rather than an error or silence. Enforce in the harness before each call, not after: check that the next call fits, so a run cannot overshoot by one expensive step. Count tokens from the provider's reported usage rather than estimates when you can.

How should permissions be designed?

Give each run the intersection of what the requesting user may do and what the task needs, enforce it in the tool layer with scoped short-lived credentials the model never sees, and gate irreversible or external actions behind approval.

Permission design for agents starts from one uncomfortable fact: the model can be steered by any text it reads. Indirect prompt injection in a document, ticket or web page can make an otherwise well-behaved agent attempt actions its user never asked for. So the question is not "will the model misuse this tool?" but "what is the worst thing this run could do if the model were fully hostile?" Permissions are how you bound that worst case.

Three principles carry most of the weight. Delegated identity: the agent acts on behalf of a specific user and should never hold more access than that user. Pass the user's identity to tools and check it there, so an agent serving a junior analyst cannot read the CFO's files. Task scoping: within the user's rights, grant only what this task needs. A report-writing run gets read access to two schemas, not the user's full write access. The effective permission is the intersection of the two. Enforcement outside the model: credentials live in the harness or tool server, are scoped and short-lived, and the tool code checks authorization on every call. The model sees tool names and schemas, never tokens.

Then tier actions by consequence. Read actions run freely within scope. Reversible writes (create a draft, add a label) run with logging. Irreversible or external writes (send email, move money, delete data, merge to main) need an approval step, a policy check, or both. Tiers can also depend on arguments: a refund under a small threshold is automatic, above it needs a human.

Watch for combinations. A run that can read private data and send data to the outside world (email, web requests, public comments) is an exfiltration path even if each tool alone is harmless. Many teams enforce a rule that a run which has read sensitive data loses its external-send tools for the rest of that run, or requires approval for them.

What is graceful degradation?

Graceful degradation means that when a model, tool or data source fails or slows down, the system falls back to a smaller but honest result, such as cached data, a simpler path or a clear partial answer, instead of failing completely or quietly guessing.

An agent depends on many things at once: one or more model providers, a retrieval index, internal APIs, a sandbox, maybe a web search service. Each fails independently. If any single failure takes down the whole response, the system's availability is roughly the product of its dependencies' availability, and five dependencies at 99.5% each give you under 98%. Graceful degradation designs, in advance, what the system does when each one is missing.

Common degradation moves: switch to a fallback model or provider when the primary returns errors or is too slow; serve cached results with their timestamp when a live data source times out; drop an optional enrichment step (a reranker, a second opinion review) and continue on the core path; reduce scope, answering the parts of a question the available tools can support; or hand off to a human queue with the work done so far. Pair these with circuit breakers: after repeated failures to a dependency, stop calling it for a cool-down period so the system does not waste its latency budget on calls that will fail.

The critical rule for LLM systems is that degradation must be honest. The worst failure mode is silent degradation into confident wrong answers. If retrieval is down and the model answers from its training data, the response looks normal but is ungrounded. If a tool times out and the agent proceeds as if it returned nothing, it may conclude a customer has no orders. The harness should mark degraded runs, and the response should tell the user what is missing ("live inventory is unavailable, figures are from 06:00").

Not every capability should degrade. For high-stakes actions, the correct degraded behaviour is often to refuse and queue: if the policy check service is down, do not skip it, block the action. Decide per dependency whether failure means fall back, reduce, or stop, and write that decision down.

How do you audit agent actions?

Keep an append-only, structured record of every consequential action: who asked, on whose behalf the agent acted, which tool ran with which arguments, what happened, which approvals and policies applied, and which model and prompt version made the decision.

Debug traces answer "why did this run behave like that?". An audit trail answers a different question, usually asked weeks later by someone outside the team: who caused this refund, deletion or email, and was it authorised? The two overlap but have different requirements. Traces can be sampled, trimmed and expired quickly. Audit records must be complete for consequential actions, tamper-evident, retained to a policy, and readable without the engineering team.

A useful audit record has a fixed schema. Actor: the human who started the task, plus the agent identity and run id. Delegation: the permission scope the run had. Action: tool name, arguments (with sensitive values redacted or hashed), target resource, timestamp. Outcome: success, failure, external reference such as the payment id. Controls: which policy rules evaluated, approval id and approver if one was required. Provenance: model, prompt version and harness version, plus a link to the full trace so an investigator can see what the model saw when it decided.

Write audit records from the harness or tool layer, never from the model. The model's description of what it did is a claim; the tool layer's record of what it executed is evidence. Make the store append-only: a separate table or log stream that the agent's credentials cannot modify, ideally with each record including a hash of the previous one so gaps or edits are detectable. Log the intent before executing a side effect and the result after, so a crash in between leaves an open intent you can reconcile rather than a silent gap.

Decide scope deliberately. Audit every write and every read of sensitive data; reads of public or low-risk data can live in normal traces. Apply the same privacy rules as any log: redact personal data that is not needed, restrict who can query the audit store, and set retention to match legal and contractual needs.

How do you verify output before action?

Put deterministic checks between what the model proposes and what gets executed: schema validation, business-rule and policy checks, a fresh authorization check, a dry run or diff where possible, and human approval for high-impact actions, with clear errors fed back on failure.

When model output drives an action, its mistakes stop being wrong words and become wrong transactions. The defence is a verification gate: a sequence of checks that every proposed action passes before execution. Order them from cheapest and most certain to most expensive, and make each one deterministic where you can.

Structural checks come first: does the output parse, match the JSON schema, use known enum values and reference ids that exist? Constrained decoding helps here, but still validate, because a schema-valid object can contain an order id that does not exist. Semantic and business-rule checks come next: the refund does not exceed the order total, the meeting is in the future, the SQL is a single read-only SELECT, the purchase order's line items sum to its total. Authorization is checked again at execution time against the acting user, not just when the run started. Policy rules, often in a policy engine, handle limits and combinations (no payments to vendors added in the last 7 days).

For actions with large or hard-to-see effects, add a preview. Run the change in dry-run mode, compute a diff, or estimate impact ("this UPDATE matches 48,112 rows"), then check the preview against expectations. A query expected to touch one customer that would touch forty thousand rows is blocked regardless of how valid it looks. For the highest-impact or least reversible actions, route the preview to a human approver who sees the exact payload, not the model's description of it.

When a check fails, return a precise error to the model ("amount 140.00 exceeds order total 92.50") so it can correct itself, and cap the number of correction attempts. Model-based review (a second model judging the action) can catch problems rules cannot express, but treat it as an extra layer, not a replacement for deterministic checks, because it shares the same vulnerabilities to persuasive or injected text.

What makes a production harness reliable?

A reliable harness survives crashes without losing or repeating work, bounds every run, contains every action, checks results against evidence rather than the model's claim, degrades honestly, and records enough to explain any outcome after the fact.

Reliability for an agent harness is not the same as the model being accurate. Models will keep making mistakes at some rate; a reliable harness makes those mistakes bounded, visible and recoverable. It helps to think in terms of the guarantees you want and the mechanism that provides each one.

No lost work, no duplicated effects. Durable state and checkpoints mean a crash or deploy resumes the run. Idempotency keys and an intent log mean resumed or retried steps do not send the email twice. Bounded runs. Budgets on steps, tokens, spend and time, plus loop detection (the same tool with the same arguments three times in a row), guarantee every run ends. Contained actions. Sandboxes, least-privilege tools and approval gates bound the worst case even when the model is wrong or manipulated. Verified completion. Success is decided by observable checks, such as tests passing, the record existing in the target system or a validator passing, not by the model saying "done".

Honest degradation. Each dependency has a defined failure behaviour, safety checks fail closed, and degraded runs are tagged and disclosed. Observability. Every model call, tool call, retry, budget event and approval is traced with a run id, and consequential actions also go to an append-only audit log. Controlled change. Prompts, tool schemas, model versions and harness code are versioned together, evaluated on a regression set and rolled out gradually, because a prompt edit can change behaviour as much as a code change.

Finally, reliability has to be measured. Track run-level metrics: completion rate, verified success rate, budget-exhaustion rate, approval rate, retries per run, cost per successful task and time to resume after failure. A harness whose success rate is unknown is not reliable, however well it is built.

What are harness hooks, and what should they enforce?

Hooks are points in the agent loop, such as before a model call, before or after a tool call, and before the run stops, where the harness runs your deterministic code to block, modify, log or check what the model is about to do.

A runtime's loop has natural seams: a run starts, the model is about to be called, the model has proposed a tool call, a tool has returned, the model says it is finished, an error occurred. A hook is a function the harness calls at one of those seams, with the relevant data, and whose return value can change what happens next. Many agent frameworks and coding-agent tools expose hooks under names like pre-tool-use, post-tool-use and stop; if yours does not, they are easy to add to a hand-written loop.

Hooks are where rules that must always hold get enforced, because they run as code on every event regardless of what the prompt says. Pre-tool hooks block or rewrite calls: deny shell commands matching dangerous patterns, refuse writes outside the working directory, require approval for production database access, strip secrets from arguments. Post-tool hooks act on results: run the linter or formatter after every file edit and feed problems back, truncate oversized output, scan results for injection markers, write the audit record. Stop hooks check completion: before the run is allowed to end, run the tests or validator, and if they fail, send the failure back and keep going (with a cap so it cannot loop forever).

The stop hook is one of the most effective reliability tools in agent harnesses. Models tend to declare success early. A stop hook converts "I have fixed the bug" into an observable check: the harness runs the test suite, and only a passing result lets the run end. The same pattern works outside coding: a report agent cannot finish until every claim has a citation that resolves.

Keep hooks fast, deterministic and narrow. A hook that calls a slow external service on every tool call adds latency to every step. A hook that fails should fail closed for safety checks (block the action) and fail open only for optional enrichment. And log every hook decision, because a blocked action is often the first sign of a prompt injection or a regression.

How do you keep side effects safe when a run is retried or resumed?

Give every intended side effect a stable idempotency key, record the intent before executing and the result after, check that log on retry or resume, and reconcile any intent whose outcome is unknown with the target system before acting again.

Retries and resumes are where agents most often do things twice. A tool call to send an invoice times out. Did the invoice go out? The harness cannot tell from a timeout. Retry blindly and the customer may get two invoices; skip the retry and they may get none. Crash recovery has the same problem at a larger scale: the worker died somewhere between deciding to act and recording that it acted.

The core tool is the idempotency key: a unique identifier for one intended action, derived deterministically from the run and step (for example run_77f3:step_6), not generated fresh on each attempt. Send it with the request; the receiving system stores it and returns the original result for any repeat. Many payment and messaging APIs support this natively. For tools that do not, the tool wrapper keeps its own table of keys and results and checks it before calling the downstream system.

Around the key, use an intent log. Before executing, write a row: key, tool, arguments, status pending. After the call, update it to succeeded with the external reference, or failed with the error. On retry or resume, the harness reads the log first. A succeeded intent is skipped and its stored result replayed to the model. A pending intent with no result is in doubt: query the target system ("does an invoice with this key or reference exist?") before deciding to resend.

A related trap is model nondeterminism. If a run resumes and re-asks the model what to do, it may propose a slightly different action, such as a different amount or wording, with a new step number and therefore a new key, and the duplicate slips through. Persist the model's proposed action along with the intent, and on resume replay the recorded proposal rather than asking again. Where you control the write path, the transactional outbox pattern helps: write the business change and the outgoing message in one database transaction, and let a separate relay deliver messages with deduplication.

How do you test the harness itself?

Test the harness separately from the model by replacing the model with scripted or recorded responses, then inject failures such as malformed tool calls, timeouts, crashes and loops, and assert that limits, permissions, retries and resume behave exactly as designed.

Model evals tell you whether the agent makes good decisions. They do not tell you whether the harness enforces a budget, blocks a forbidden tool, or resumes without duplicate sends, and real model behaviour is too variable to trigger those paths reliably. Harness tests need the model out of the loop, so the harness's own logic becomes deterministic and testable like any other code.

The main technique is a scripted fake model: an object with the same interface as the model client that returns a predetermined sequence of responses. One script calls a tool that does not exist. Another emits malformed JSON arguments. Another requests the same search forever. Another calls a write tool outside its permissions, as an injected model would. Each test asserts the harness's response: error fed back, call blocked, run stopped at the budget with status budget_exhausted, audit record written.

Second, fault injection on everything else. Wrap tools and the provider client so tests can make them time out, return 429 or 503, return huge outputs, or raise halfway through. Assert that retries follow the policy, that non-retryable errors fail fast, that output is truncated before it enters context, and that degraded runs are tagged. Then test crash and resume: run the harness in a subprocess, kill it at random points, restart it, and check invariants such as each email sent exactly once and the final state matching an uninterrupted run.

Third, record and replay real runs. Capture model responses and tool results from production traces (with personal data removed) and replay them through a new harness version. Because the model's outputs are fixed, any difference in tool execution, policy decisions or final status is caused by the harness change. This catches regressions such as a refactor that stops writing audit records for one tool. Keep these tests fast and run them in CI on every harness change; reserve live-model evals for measuring decision quality.