The GenAI Field Guide

Evals and testing

Measure quality with realistic cases before and after every change.

An eval is how a team finds out whether an AI system is good enough, and whether a change made it better or worse. Traditional software can be tested with assertions because the correct output is usually one known value. A language model produces open-ended text, takes different paths through tools, and can give a different answer to the same input on a second run. Quality has several dimensions at once: correctness, groundedness, tone, format, safety, cost and latency. Evals turn those fuzzy judgements into repeatable numbers you can track, compare and gate releases on.

The mental model is a measurement instrument with three parts: a dataset of realistic cases, a grader that scores each output (code checks, model judges or people), and a harness that runs the system over the dataset and aggregates the scores. Every part can be wrong. A dataset that only holds easy questions flatters the system, a judge that prefers long answers rewards padding, and a sample of 30 cases cannot tell a 3-point improvement from noise. Most of the craft is in making the instrument trustworthy before trusting what it says.

The questions build in that order. The Basic questions cover what an eval is, how to build datasets and golden sets, and the main grading methods: exact match, model judges and groundedness checks. The Advanced questions cover when to trust a judge, how to evaluate agents and RAG pipelines, how many cases you need, how to handle randomness, and how evals continue into production through shadow tests, canary releases and red teaming. Read it as one loop: find failures, turn them into cases, grade them reliably, and gate every change on the result.

What is an LLM eval?

An LLM eval is a repeatable test that runs an AI system over a fixed set of inputs, scores each output against defined criteria, and reports aggregate quality so you can compare versions.

An eval has three parts. A dataset holds the inputs, plus expected answers or grading rules. A grader scores each output: a code check, a model acting as judge, or a human reviewer. A harness runs the system over every case, collects the scores and summarises them, for example "87% of answers correct, 2% cited a wrong policy, median latency 1.4 seconds".

The point is repeatability. A demo shows that a system can work once; an eval shows how often it works across the cases you care about. Because the dataset and graders stay fixed, you can run the same eval after a prompt edit, a model upgrade or a retrieval change, and see whether quality moved. Without that, teams judge changes by trying a handful of inputs by hand, which overweights whatever the tester happened to type.

Evals measure the system, not just the model. A support bot's quality depends on its prompt, retrieval, tools, model and post-processing together. Public benchmarks tell you about a model on someone else's task; your own eval tells you about your system on your task, which is the number that matters for a release decision.

Good evals score several dimensions separately rather than one blended number. An answer can be correct but rude, polite but ungrounded, or right but too slow. Keeping the dimensions apart tells you what to fix. The trade-off is effort: a useful eval needs realistic cases and graders you trust, and both take real work to build and maintain.

Why are unit tests alone insufficient?

Unit tests assert one exact expected result, but LLM outputs vary in wording, can differ between runs, and have several quality dimensions at once, so pass or fail on exact output misses most of what matters.

A unit test encodes a single correct answer: assert add(2, 2) == 4. That works when the function is deterministic and the correct output is unique. A summarisation feature has neither property. Two summaries can use completely different sentences and both be excellent, and the same prompt can produce different text on a second call. An exact-match assertion would fail good outputs and, if loosened enough to pass them, would also pass bad ones.

Quality is also multi-dimensional. A reply can be factually right but cite the wrong source, follow the format but drop a required caveat, or be correct but take 12 seconds. Unit tests check one property at a time against one value; evals score many properties over many inputs and report rates, such as "94% grounded, 3% wrong tone".

Unit tests still matter. Everything deterministic around the model should be tested the normal way: prompt templating, output parsing, schema validation, tool argument handling, retry logic and permission checks. Some model properties can also be asserted directly: the output is valid JSON, a field is one of five allowed values, no email address appears in the reply. The rule is to use code checks for properties that code can decide, and evals for judgements over distributions.

The trade-off is speed and cost. Unit tests run in milliseconds with no API calls; a full eval may take minutes and cost money per run. Many teams therefore run unit tests on every commit, a small fast eval subset on every pull request, and the full eval before release.

What is an eval dataset?

An eval dataset is a versioned collection of representative inputs, each paired with an expected output, required facts or grading rules, chosen to cover normal traffic, hard cases and known failure modes.

Each case is an input plus whatever the grader needs to decide whether the output is good. That might be an exact label ("refund"), a list of facts the answer must mention, a reference answer, or a rubric ("must decline and point to the human agent"). Cases usually carry metadata too: a stable id, the source, the category, and a difficulty or risk tag, so you can slice results later.

The dataset should look like the traffic the system will face, not like the examples that were easy to write. Good sources are production logs (with personal data removed), support tickets, user research, and failures found during error analysis. Synthetic cases generated by a model help fill gaps, such as rare languages or edge cases, but they tend to be cleaner and more uniform than real input, so they should be a minority and reviewed by a person.

Coverage matters more than size at first. A useful starting mix is mostly typical cases, a meaningful slice of hard or ambiguous ones, adversarial inputs, and cases where the right behaviour is to refuse or ask a question. Tagging the slices lets you see that the system scores 95% on typical questions but 60% on multi-part ones, which a single average hides.

Treat the dataset like code: version it, review changes, and never edit a case silently because the system fails it. Keep a held-out portion that prompt authors do not look at, otherwise the prompt gets tuned to the visible cases and the score stops predicting real performance.

What is a golden dataset?

A golden dataset is a smaller, carefully reviewed set of eval cases whose expected answers were checked by domain experts, so it can serve as the trusted reference for scoring systems and calibrating automated judges.

Every eval case has an expected outcome, but not every expected outcome is equally trustworthy. Labels written quickly by an engineer, or generated by a model, contain errors. A golden set is the subset where qualified people have reviewed each answer, resolved disagreements, and recorded why. When the golden label and the system disagree, you can be confident the system is wrong.

Golden sets do two jobs. First, they are the release benchmark: the number you report to stakeholders comes from cases you trust. Second, they calibrate other graders. Before relying on a model judge, you run it on the golden set and measure how often it agrees with the experts. If agreement is poor, the judge is not ready.

Building one is expensive, which is why golden sets are usually hundreds of cases, not thousands. A common process is to have two experts label each case independently, measure their agreement, and send disagreements to a third reviewer or a discussion. Low agreement between experts is itself useful: it often means the task or rubric is ambiguous and needs a clearer definition before any automated grading can work.

Golden sets go stale. Policies change, products launch, and user behaviour shifts. Schedule periodic review, mark cases with the date and policy version they reflect, and retire cases whose correct answer has changed rather than leaving them to punish a correct system.

What is exact-match evaluation?

Exact-match evaluation checks whether the output equals an expected value, usually after normalising case, whitespace and formatting. It is cheap, deterministic and reliable for labels, codes and extracted fields, but useless for free text.

Exact match compares the model's answer with a known correct value and returns pass or fail. It suits tasks with one right answer: classification labels, routing decisions, yes or no questions, extracted dates, IDs and amounts. Because it needs no judge and no human, it is fast, free to run and gives the same result every time, which makes it the backbone of many eval suites.

Raw string equality is usually too strict. "Approved", "approved." and " APPROVED" are the same answer. Practical exact match normalises first: lowercase, trim whitespace, strip punctuation, map synonyms, and parse numbers and dates into a canonical form. For structured output, parse the JSON and compare field by field, so key order and spacing do not matter and you can report which fields were wrong.

Several close relatives extend the idea. Set match compares unordered lists, often with precision and recall. Numeric tolerance accepts values within a range. Regex or contains checks that a required pattern appears. Execution match runs generated code or SQL and compares the result instead of the text, which accepts different but equivalent queries.

The limitation is clear: exact match cannot judge a paragraph. If many phrasings are correct, it will fail good answers. The design move is to make more of the task exact-matchable, for example by asking the model to return a structured label alongside its explanation, then grading the label with code and the explanation separately.

What is an LLM-as-a-judge?

LLM-as-a-judge means using a language model, given a rubric, to grade another system's output on qualities code cannot check, such as helpfulness, tone or whether an answer cites its evidence.

Many quality questions have no exact answer: is this reply polite, does it address every part of the question, does it make claims the sources do not support? Human review answers them well but does not scale to thousands of cases per change. A judge model reads the input, the output and a rubric, and returns a score or label with a short reason. It can grade a full eval run in minutes.

Judges work best on narrow, well-defined questions. "Rate this answer from 1 to 10" produces vague, drifting scores. "Does the answer state the refund window from the policy? Answer PASS or FAIL" produces stable, checkable ones. Binary or small categorical scales are easier to calibrate than wide numeric ones, and one criterion per judge call is more reliable than a combined rubric.

Give the judge what it needs to decide. For correctness, include a reference answer or the source documents. Ask for the reasoning before the verdict so the decision is grounded in the reasoning, and request structured output so you can parse the result reliably. Keep the judge prompt versioned like any other prompt, because changing it changes every score.

The trade-off is that a judge is itself a model with errors and biases. It may prefer longer answers, the first answer it sees, or text written in its own style. A judge is a measuring instrument, and its scores mean little until you have checked how often it agrees with people on a labelled sample.

What is a regression eval?

A regression eval reruns a fixed set of cases the system previously handled correctly, after every change, to catch anything the change broke. It turns past fixes and known failures into a permanent safety net.

LLM systems are unusually prone to regressions because their parts interact in ways that are hard to predict. Adding a sentence to a prompt to fix refund questions can change how the model handles shipping questions. Upgrading the model can improve reasoning and break a format the parser relied on. A regression eval answers one question after each change: did anything that used to work stop working?

The suite grows from experience. Each time a bug is found, in testing or production, the failing input and the correct behaviour become a new case. Over months the suite becomes a record of every way the system has failed, and every change is checked against all of them. This is the same discipline as regression tests in normal software, applied to probabilistic behaviour.

Regression evals usually run as a gate in continuous integration. The harness compares the new run with a stored baseline and fails the build if the pass rate drops beyond a threshold, or if any case tagged critical fails. Report the specific cases that flipped from pass to fail, not just the score, because a stable average can hide five newly broken cases offset by five newly fixed ones.

Because outputs vary between runs, a small number of flips may be noise. Running flaky cases several times, or setting a tolerance of one or two cases on the non-critical set, keeps the gate useful without blocking every change. Critical cases, such as safety or compliance behaviour, should have zero tolerance.

What is groundedness?

Groundedness, also called faithfulness, measures whether every claim in an answer is supported by the evidence the system was given, such as retrieved documents or tool results, regardless of whether the claim happens to be true.

Groundedness asks a narrow question: can each statement in the answer be traced to the supplied sources? It is different from correctness. An answer can be true but ungrounded, when the model fills a gap from its training data, and grounded but wrong, when the source document itself is out of date. In systems built on retrieval, groundedness is often the property you most need, because it means answers come from your approved content rather than from the model's memory.

The standard method is to break the answer into atomic claims, then check each claim against the sources. "Refunds take 5 to 7 days and require a receipt" becomes two claims. A verifier, usually a judge model or a natural language inference model trained to decide whether one text entails another, labels each claim as supported, contradicted or not found. The groundedness score is the share of supported claims, and any contradicted claim can be flagged as severe.

Groundedness depends on retrieval quality. If the right passage was never retrieved, a grounded answer is impossible, and the right behaviour is to say so. That is why RAG evaluation scores retrieval and generation separately: low groundedness with good retrieval points at the prompt or model; low groundedness with poor retrieval points at search.

The trade-off is cost and strictness. Claim-level checking needs several model calls per answer, and verifiers can be too strict about reasonable inferences or too lenient about paraphrased numbers. Calibrate on human-labelled examples, and decide in the rubric how to treat common knowledge such as "Monday comes after Sunday".

Can an LLM judge be trusted?

Only after calibration. A judge has measurable biases, including position, length and self-preference, so you trust it to the degree that it agrees with human experts on a labelled sample, and you keep checking that agreement over time.

A judge is a measurement instrument, and instruments need calibration. Research on model judges, including the 2023 paper Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, found that strong judges can agree with human preferences about as often as humans agree with each other, but also documented systematic biases. Position bias: in pairwise comparisons the judge favours whichever answer appears first (or second). Verbosity bias: longer, more detailed answers score higher even when they are not better. Self-preference: a judge tends to rate outputs from its own model family more favourably. Judges also struggle to grade answers to questions they cannot solve themselves, such as hard maths.

Calibration means building a human-labelled sample, typically 100 to 300 cases drawn from real outputs, and measuring agreement. Raw agreement is a start, but use a chance-corrected statistic such as Cohen's kappa, and look at the confusion matrix. A judge that agrees 90% of the time can still be useless if 90% of cases pass and it simply says PASS to everything. What you usually care about is how many real failures it catches (recall on FAIL) and how often its FAILs are real (precision).

Many biases can be reduced by design. Narrow binary criteria are more reliable than holistic scores. Reference answers or source documents let the judge check facts instead of guessing. Swapping the order of candidates and counting only consistent verdicts cancels position bias. Asking for reasoning before the verdict helps on multi-step criteria. Using a judge from a different model family than the system under test reduces self-preference.

Trust also decays. A judge calibrated on last quarter's outputs may drift when the system changes style or the judge model is updated by its provider. Pin the judge model version where the provider allows it, rerun calibration when anything in the judge changes, and keep a small weekly human review of judge decisions. When judge and human disagree, read the cases; disagreements often reveal an ambiguous rubric rather than a bad judge.

What is trajectory evaluation?

Trajectory evaluation scores the sequence of steps an agent took, including tool calls, arguments, order and cost, rather than only its final answer, because an agent can reach the right answer through an unsafe, wasteful or lucky path.

An agent's output is a trajectory: the list of model turns, tool calls, tool results and decisions that led to the final answer. Final-answer evals miss important failures. An agent that issues a refund before verifying the order, calls the search tool 14 times when 2 would do, or reads a customer's records it had no reason to access may still produce a correct final message. In production, the path is often where the risk and cost live.

Trajectory checks fall into a few families. Required steps: the agent must call verify_identity before issue_refund. Forbidden steps: no write tool in a read-only task, no call to a tool outside its allowlist. Argument correctness: the refund amount matches the order total. Efficiency: number of steps, tokens and latency under a budget. Recovery: after a tool error, the agent retries sensibly or stops rather than looping. Most of these are deterministic checks over the logged trace and need no model judge.

Avoid scoring against one exact reference path. Agents can legitimately solve a task in different orders, and penalising any deviation punishes valid strategies. Prefer partial-order constraints (A before B), set membership (these tools must appear, those must not) and budgets. For softer questions, such as whether a step was reasonable given what the agent knew, a judge can read the trajectory, but calibrate it like any other judge.

Trajectory evaluation depends on good tracing. Every tool call, argument and result needs to be logged with step order, so the same records serve production debugging and offline evals. For tools with side effects, run evals against sandboxed or mocked tools, and include the mocked responses in the case so runs are reproducible.

What is pairwise evaluation?

Pairwise evaluation shows a judge two outputs for the same input and asks which is better. It is more sensitive than absolute scoring for comparing versions, but needs order swapping and tie handling to be reliable.

Absolute scoring asks "rate this answer 1 to 5"; pairwise asks "which of these two answers is better?". People and model judges are both more consistent at relative comparisons than at absolute ratings, because the second answer provides an anchor. When you are deciding between an old and a new prompt, pairwise judging often reveals a difference that absolute scores blur, since both versions cluster at 4 out of 5.

The standard protocol runs both systems on the same inputs, then shows each pair to a judge blind: labelled A and B, with no hint which is new. Because of position bias, run every pair twice with the order swapped. If the judge picks the same underlying answer both times, count it; if it flips, count a tie. Report win, loss and tie rates, for example "new wins 46%, old wins 31%, tie 23%", and test whether the difference is larger than chance with a sign test over the decisive pairs.

Pairwise results can be aggregated across many systems into a ranking with models such as Bradley-Terry or Elo ratings, which is how public preference leaderboards are built. For product work, the two-way comparison is usually enough.

The trade-offs are real. Pairwise tells you which is better, not whether either is good enough: two bad answers still produce a winner. It costs two judge calls per pair, and it does not compose: you cannot compare version C with A using only A-B and B-C results without assuming the preferences are consistent. Use pairwise to choose between candidates, and absolute checks (exact match, groundedness, safety) to confirm the winner meets the bar.

What is shadow testing?

Shadow testing runs a candidate system on real production inputs alongside the live one, logs both outputs, and never shows the candidate's output to users, so you can compare behaviour on true traffic without user risk.

Offline eval sets are always a sample, and real traffic contains phrasings, languages and edge cases nobody thought to include. Shadow testing closes that gap. Each incoming request goes to the production system as normal, and a copy goes to the candidate, usually asynchronously so it adds no latency. The user only sees the production response. Both outputs are logged with the same request id for comparison.

The comparison uses the same graders as offline evals: code checks for format and policy, judges for quality, pairwise comparison between live and candidate, and operational metrics such as latency, token use and error rate. Shadow traffic is also the best way to find disagreements: requests where the two systems behave very differently. Reviewing a sample of those by hand is often more informative than any aggregate score.

Side effects are the main hazard. A shadow agent that can call write tools might send emails or issue refunds twice. Shadow systems must run with write tools disabled, stubbed or pointed at a sandbox, which means shadow testing measures what the candidate would have done, not the full end-to-end outcome. Multi-turn conversations add another complication: the shadow sees the conversation history produced by the live system, not its own, so later turns are not a faithful test of the candidate.

Shadowing costs real money because every shadowed request is paid for twice, and it yields no user feedback such as clicks or ratings. Teams usually shadow a sample, for example 10% of traffic for one to two weeks, rather than everything. It also needs the same privacy controls as production, since the candidate sees real user data.

What is canary deployment?

Canary deployment releases a change to a small share of real users first, watches quality, safety, latency and cost against the current version, and widens the rollout only if every guardrail metric holds.

A canary is the first live exposure of a change to users. Unlike shadow testing, users actually see the canary's output, so you also measure what they do: ratings, follow-up questions, escalations to humans, task completion and complaints. Because only a small share of traffic is exposed, a bad change affects few users and can be rolled back quickly.

Routing should be sticky: assign each user or conversation to one version for the whole session, typically by hashing a user id, so a conversation does not switch models mid-thread and so per-user metrics stay clean. Traffic then increases in stages, for example 1%, 5%, 25%, 100%, with a minimum observation window at each stage and explicit criteria for moving on.

Decide the guardrail metrics and thresholds before the rollout. Typical ones are error and timeout rate, p95 latency, cost per request, safety-filter triggers, judge-scored quality on sampled traffic, escalation rate and thumbs-down rate. Automatic rollback on hard limits (errors, safety) avoids waiting for a person to notice. Softer quality signals usually need a human decision, because they move slowly and have wide error bars at small traffic shares.

Canaries have statistical limits. At 1% of traffic, a rare failure that happens once per 10,000 requests may not appear for days, and a 2-point change in thumbs-down rate may be undetectable. That is why canaries complement offline evals and shadow tests rather than replacing them: offline evals catch known failure modes before exposure, and the canary catches what only real users reveal.

How many examples are enough?

Enough that the uncertainty in your score is smaller than the difference you need to detect. Around 100 cases detects only large changes; spotting a few points of difference needs several hundred paired cases, plus coverage of every important slice.

An eval score is an estimate from a sample, so it has an error bar. For a pass rate, the 95% confidence interval is roughly plus or minus 1.96 * sqrt(p * (1 - p) / n). With 100 cases and an 80% pass rate, that is about plus or minus 8 points. A move from 80% to 84% on 100 cases is well inside the noise. With 400 cases the interval shrinks to about plus or minus 4 points; with 1,000, about 2.5. Halving the error bar needs four times the cases.

Comparisons are more efficient than two independent scores. When both versions run on the same cases, most cases pass or fail for both, and only the cases that flip carry information. Tests designed for paired results, such as McNemar's test or a paired bootstrap that resamples cases, detect smaller differences than comparing two separate confidence intervals. This is one reason to keep the eval set fixed across versions.

Size is only half the question; coverage is the other. Five hundred near-identical easy questions give a precise estimate of the wrong thing. Make sure every slice that matters, such as each product area, language, refusal case and known failure mode, has enough cases to read on its own; a slice with 8 cases can only show catastrophic changes. Critical behaviours like safety refusals may need dedicated suites where even one failure blocks release, regardless of the overall rate.

A practical path is to start with 50 to 100 diverse, real cases so you have something to iterate against in week one, grow to a few hundred as error analysis surfaces new failure types, and add a larger held-out set for release decisions. Cost and latency of running the eval set the upper bound: if the suite takes an hour, people stop running it, so keep a fast subset for every change and a full set for releases.

What is red teaming?

Red teaming is deliberate adversarial testing: people and automated attackers try to make the system leak data, take harmful actions, produce unsafe content or ignore its instructions, so weaknesses are found and fixed before real attackers find them.

Normal evals test whether the system works for cooperative users. Red teaming tests what happens when someone is trying to break it. For LLM systems the attack surface is unusually broad: direct prompt injection and jailbreaks in the user message, indirect injection hidden in retrieved documents, emails or web pages the system reads, attempts to extract the system prompt or other users' data, misuse of tools ("refund every order on this account"), and harmful or off-policy content. Agents with tools raise the stakes, because a successful injection can cause actions, not just text.

Effective red teaming combines approaches. Manual testing by people who understand the domain finds creative, context-specific attacks. Automated red teaming uses an attacker model or mutation scripts to generate thousands of variants of known attacks and score whether they succeed. Structured catalogues, such as the OWASP Top 10 for LLM Applications and MITRE ATLAS, help make sure common categories are covered. Each attack needs a clear success criterion: data leaked, forbidden tool called, policy-violating text produced.

Findings should feed back in two ways. First, fix the root cause with controls that do not depend on the model obeying: least-privilege tools, permission checks outside the model, output filters, and confirmation for risky actions. Second, turn every successful attack into a regression case so the fix stays fixed. A prompt instruction like "ignore malicious instructions" can help, but it is not a control, and red teams usually find a way around it.

Red teaming is never complete. New attack techniques appear, models change, and every new tool or data source adds surface. Run a focused exercise before each significant launch or capability change, keep an automated adversarial suite in CI, and track attack success rate over time rather than treating one clean exercise as proof of safety.

How do you evaluate a RAG system?

Evaluate retrieval and generation separately: measure whether the right passages were retrieved, then whether the answer is grounded in them, correct and complete. Separate scores tell you which half of the pipeline to fix.

A retrieval-augmented generation (RAG) system has two stages that fail differently. Retrieval can miss the passage that holds the answer, rank it too low, or return stale or forbidden content. Generation can ignore good passages, add unsupported claims, or answer when it should say the documents do not cover the question. A single end-to-end score cannot tell you which stage failed, so most teams score each stage on its own and also score the whole.

For retrieval, label which documents or passages are relevant for each question, then compute metrics such as recall@K (did any relevant passage appear in the top K?) and ranking measures such as mean reciprocal rank. Recall matters most, because the model cannot use what it never saw. Labels can come from experts, or from a judge that rates each retrieved passage's relevance, calibrated on a human sample.

For generation, score groundedness (are the claims supported by the retrieved passages?), correctness against a reference answer, completeness (did it cover every part of the question?), citation accuracy (does each citation point to a passage that supports it?) and appropriate abstention. Include questions whose answers are not in the corpus, because a system that never says "I don't know" will look fine on answerable questions and invent answers to the rest.

Reading the two scores together gives a diagnosis. High recall with low groundedness points to the prompt, the model or too much noisy context. Low recall points to chunking, embeddings, query rewriting or missing documents. Also test retrieval against permission rules (a user should never retrieve documents they cannot read), and re-run the eval when the index changes, since adding documents can change rankings for old questions.

What is error analysis, and why do it before choosing metrics?

Error analysis means reading a sample of real outputs, labelling what went wrong in each, and grouping the failures into categories. It tells you which evals to build, instead of measuring generic qualities that may not be your real problems.

Teams often start by picking off-the-shelf metrics such as helpfulness, coherence or toxicity, then wonder why the scores look fine while users complain. The metrics were chosen before anyone knew how the system actually fails. Error analysis reverses the order. You read real outputs first, write a short free-text note on each failure, and only then decide what to measure.

The process borrows from qualitative research. In the first pass, a person with domain knowledge reads 50 to 100 traces and writes open notes: "quoted the 2024 policy", "ignored the second question", "called search three times with the same query". In the second pass, the notes are grouped into a small set of failure categories and counted. The counts give priorities: if 40% of failures are stale policy citations and 3% are tone, the first eval and the first fix are obvious.

Each important category then becomes a targeted eval: a code check where possible ("answer cites a document older than the current policy version"), a narrow judge where not, and a set of cases drawn from the examples you found. Because the judge criteria come from real failures, they are specific enough to calibrate, and the cases are realistic by construction.

Error analysis is not a one-off. Repeat it after major changes and on a regular sample of production traffic, because fixing the top failure reveals the next one, and new features create new failure types. It feels slow compared with running an automated metric, but an hour of reading traces routinely finds problems no existing metric was looking for.

How do you handle nondeterminism in evals?

Run each case several times, report pass rates per case rather than one sample, and measure consistency explicitly. Temperature zero reduces variation but does not guarantee identical outputs, so evals must be designed for variance.

The same input can produce different outputs across runs. Sampling with temperature above zero is the obvious cause, but even at temperature zero many hosted models are not perfectly deterministic: batching, parallel floating-point arithmetic and infrastructure changes can alter results, and providers may update a model behind the same name. Agents amplify the effect, because one different early step sends the whole trajectory down another path.

The practical response is to treat each case's result as a probability. Run each case k times (often 3 to 5 for regression suites, more for critical cases) and record the pass rate per case. This separates three groups: cases that always pass, cases that always fail (real bugs), and flaky cases that pass sometimes. Flaky cases are often the most informative, because they mark where the system is on the edge of its ability or where the prompt is ambiguous.

Two summary metrics answer different questions. pass@k asks whether at least one of k attempts succeeded, which suits workflows with a verifier that can pick the good attempt, such as code that must pass tests. pass^k (pass to the power k) asks whether all k attempts succeeded, which matters for user-facing agents where every customer gets one try. A system with 80% per-run success has pass^3 of only about 51% if runs are independent, which is why agents that look good in demos frustrate users.

Control what you can. Use temperature zero for judges, pin model versions where the provider allows it, fix seeds where supported, mock tool responses so the environment does not vary, and log every run's full configuration. Then design gates for variance: compare per-case pass rates between versions rather than single runs, and only flag a regression when a case's rate drops by more than its observed run-to-run spread.