The GenAI Field Guide

AI safety and reliability

Design systems that can admit uncertainty, fall back, and prevent harmful actions.

A language model is a probabilistic component. It will sometimes be wrong, sometimes be confidently wrong, and sometimes be steered by text it should have ignored. Reliability engineering for GenAI accepts that and asks a different question: when the model is wrong, what is the worst thing that can happen, and how do we make that outcome small, visible and recoverable? The answer is rarely a better prompt. It is a set of deterministic checks, limits and fallbacks around the model, owned by code the model cannot talk its way past.

The mental model is defence in depth around three boundaries. An input guardrail decides what reaches the model's context. An output guardrail decides what the model's text is allowed to become: a message to a user, a parsed record, a citation. An action guardrail decides whether a proposed tool call is allowed to change the world. Each layer is imperfect, so each assumes the one before it failed. Behind the layers sit the behaviours a reliable system needs when the checks say no: abstention (declining to answer), graceful failure (failing in a small, honest way) and fallback (a planned alternative path).

The Basic questions define these layers and behaviours. The Advanced questions deal with the hard parts: whether the model's own confidence can be used to decide when to abstain, how to measure that with calibration, how to validate and gate actions, which actions need a human, how to bound the damage of a mistake you did not foresee, and how a policy engine turns all of this into rules you can test and audit. The final questions cover the kill switch for when things go wrong at scale and how to tell whether your guardrails actually work.

What are guardrails?

Guardrails are checks and limits placed around a model that decide what goes into it, what its output may become, and which of its proposed actions may run. They live in code outside the model, so a persuasive or injected prompt cannot switch them off.

A model follows instructions most of the time, but "most" is not a safety property. Guardrails exist because you cannot make a probabilistic component deterministic by asking nicely. A guardrail is any check or constraint that sits on the path between the model and something you care about: the model's context, the user's screen, a downstream system, or the outside world.

It helps to group guardrails by where they sit. Input guardrails run before text reaches the model: redacting secrets, blocking disallowed requests, flagging likely prompt injection in a retrieved document. Output guardrails run on what the model produced: schema validation, citation checks, toxicity or PII filters. Action guardrails run before a tool call executes: permission checks, amount limits, human approval. Each layer catches failures the others miss, which is why serious systems use all three.

Guardrails come in two strengths. Deterministic ones (regular expressions, schema validators, allowlists, permission lookups) are fast, predictable and cheap, but they only catch what you anticipated. Model-based ones (a classifier or a second model acting as a judge) catch fuzzier problems such as an off-topic answer or a subtle insult, but they add latency, cost and their own error rate. A common pattern is cheap deterministic checks on everything and model-based checks only where the risk justifies them.

The key design rule: the guardrail must not be the model policing itself in the same prompt. "Never reveal customer data" in a system prompt reduces how often the model does it; a filter that strips account numbers from the output prevents it. The trade-off with every guardrail is false positives (blocking good requests, which annoys users) against false negatives (letting bad ones through). Choose the threshold from the cost of each, not from a vendor default.

What is an input guardrail?

An input guardrail is a check that runs before text reaches the model or enters its context. It redacts secrets and personal data, rejects requests the product should not serve, and marks untrusted content such as retrieved documents so it is treated as data, not instructions.

Everything that enters a model's context window can influence its output, and much of it also ends up in logs, traces and sometimes provider-side storage. An input guardrail decides what is allowed in and in what form. It runs on every source of text: the user's message, uploaded files, retrieved documents, tool results and messages from other agents.

Common input checks fall into four groups. Redaction replaces secrets, card numbers, national IDs or other PII with placeholders, either because the model does not need them or because policy forbids sending them to a third party. Scope filtering rejects requests the product is not meant to handle, such as a banking assistant being asked for medical advice. Size and format limits cap length and reject malformed or unexpected file types. Injection screening looks for text that tries to give the model new instructions, especially in content the user did not write.

Injection screening is the weakest of the four, and it is worth being honest about that. Classifiers and heuristics catch obvious phrases such as "ignore previous instructions", but an attacker can rephrase indefinitely, and natural language has no reliable boundary between data and command. Treat injection detection as a tripwire that raises a signal, not a wall. The real protection against injection is downstream: limiting what the model can do with any tool call it is tricked into proposing.

Input guardrails also shape context. Wrapping retrieved text in clearly labelled delimiters and telling the model it is untrusted reduces, though does not eliminate, the chance it follows embedded instructions. Keeping a reversible mapping for redacted values lets you restore them in the final output when the user is allowed to see them.

What is an output guardrail?

An output guardrail checks what the model produced before it is shown to a user or used by another system. It validates format, verifies claims against sources, and filters content that is harmful, off-policy or leaks data, then blocks, repairs or retries.

The model's output is untrusted until checked, in the same way user input is untrusted in a web application. An output guardrail is the point where you decide whether this specific response is fit for its destination. The destination matters: text shown to a customer needs different checks from JSON that will be written to a database.

There are three broad kinds of output check. Structural checks confirm the output parses and matches a schema, that required fields are present, that enums hold allowed values. Content checks look for things that must never appear: personal data belonging to someone else, profanity, competitor promises, regulated advice. Grounding checks verify that claims are supported, for example that every citation points to a document that was actually retrieved and that the quoted span really appears in it.

When a check fails, you have three options. Block and show a safe message, which is right for leaks or harmful content. Repair deterministically, such as stripping a disallowed field or truncating, which is right for minor format faults. Retry with the error fed back to the model, which works well for schema failures but needs a cap, usually one or two attempts, so latency and cost stay bounded.

Streaming complicates this. If you stream tokens to the user as they are generated, a check that needs the full answer runs too late. Options are to buffer short answers, to check sentence by sentence and cut the stream when something fails, or to stream only for low-risk surfaces. Pick deliberately; many teams discover the issue only after a filtered answer has already been displayed.

What is an action guardrail?

An action guardrail is a check that runs after the model proposes a tool call and before the tool executes. It decides whether this specific action, with these arguments, for this user, right now, is allowed, needs approval, or must be refused.

When a model only produces text, the worst case is a bad message. When it can call tools, the worst case is a bad action: money sent, records deleted, an email to every customer. The model never executes a tool itself. It emits a request, and your code runs it. That handoff is the single most important control point in an agent, and the action guardrail lives there.

An action guardrail checks things the model cannot be trusted to check. Authorisation: is the human the agent is acting for allowed to do this to this resource? Arguments: are the values valid, in range, and consistent with what the user actually asked? Business rules: is the refund within policy, is the account not frozen, is it within business hours? Risk tier: does this action need a human to confirm it? The decision is allow, deny, or escalate, and the reason is logged.

The reason this must be code, not prompt instructions, is prompt injection. An attacker who can get text into the context (a web page, an email, a support ticket) can make the model propose any call its tools allow. If the only protection is "only refund the current customer's orders" in the system prompt, the attacker controls the outcome. If a function compares the order's owner with the authenticated user before executing, the injected call fails no matter how persuasive the text was.

Good action guardrails also return useful feedback. A denial with a short machine-readable reason ("amount exceeds per-transaction limit of 500") lets the model explain the situation to the user or choose a smaller action, instead of retrying the same call blindly.

What is abstention?

Abstention is the system deliberately declining to answer or act when the evidence is not good enough, and saying so plainly. It trades some coverage for accuracy: fewer questions answered, but far fewer confident wrong answers.

A language model is trained to produce plausible continuations, and a plausible continuation almost always exists. Left alone, it will answer questions it has no basis for, which is where many hallucinations come from. Abstention is the design choice to make "I can't answer that reliably" a first-class outcome, decided by the system rather than left to the model's mood.

Abstention should be triggered by signals you can check. In a retrieval system, the strongest is evidence: no document scored above the relevance threshold, or the retrieved passages do not contain the requested fact. Other signals include the output guardrail failing after a retry, low agreement between several sampled answers, the request being out of scope, or a required tool being unavailable. Asking the model to say "I don't know" helps a little, but models are generally better at following that instruction when the context is obviously empty than when it is subtly insufficient.

The useful way to think about abstention is coverage versus accuracy. If you answer every question, coverage is 100% and accuracy is whatever the model achieves. If you abstain on the least certain 15%, accuracy on the remaining 85% usually rises. Where to set the threshold depends on the cost of each outcome. For a recipe bot a wrong answer is cheap; for a dosage question or an invoice approval it is expensive, so you accept lower coverage.

A good abstention is helpful, not a dead end. Say what is missing, offer what can be said with confidence, and route to a next step: a human, a document, a narrower question. "I can't verify that invoice because the PO number does not match any open order; here is the closest match" is far more useful than a refusal.

What is graceful failure?

Graceful failure means that when something breaks, the system fails in a small, clear and safe way: it tells the user what happened, keeps their work, does nothing irreversible, and avoids guessing to cover the gap.

AI systems have more ways to fail than ordinary software. The model provider can time out or rate-limit. Retrieval can return nothing. A tool can error halfway through a multi-step task. The output can fail validation twice. A graceful system has decided in advance what happens in each case. An ungraceful one either crashes with a stack trace or, worse, lets the model improvise an answer when the data it needed never arrived.

That second failure mode is specific to GenAI and is the one to design against. If a tool call to the order system fails and the error message is fed back into context, a model will often produce a fluent reply anyway, sometimes inventing the order status. Graceful failure means that a failed dependency produces an explicit failure state the model cannot paper over: the harness, not the model, decides to show "I can't reach the order system right now."

The core decision for each component is fail closed or fail open. Fail closed means that if the check or dependency is unavailable, the action does not happen. Fail open means the system proceeds without it. Safety and authorisation checks should almost always fail closed. Non-critical enrichments, such as a personalisation lookup or a tone classifier on low-risk chat, can fail open so that the core experience survives.

Graceful failure also covers partial progress. An agent that has completed three of five steps should record which ones succeeded, avoid repeating side effects on retry, and tell the user exactly where it stopped. That requires idempotent tools and persisted state, which is why graceful failure is an architecture property, not an error message.

What is a fallback?

A fallback is a planned alternative path the system takes when the preferred one fails or is not good enough: another model or provider, a simpler method, a cached or templated answer, or a hand-off to a human.

Graceful failure is about failing well; a fallback is about still succeeding, in a reduced form. The key word is planned. A fallback is chosen, tested and monitored ahead of time, with a clear trigger, rather than improvised during an incident.

Fallbacks come in a few families. Provider or model fallback routes to a second model when the primary times out, rate-limits or is down. Capability fallback switches to a simpler, more deterministic path, such as keyword search when vector search fails, or a fixed form when the conversational flow cannot extract the fields. Content fallback serves a cached or templated answer for common questions. Human fallback hands the conversation to a person, with the context so far, when the system cannot complete the task or the user asks for it.

Triggers should be specific. Infrastructure triggers include timeouts, error codes and a circuit breaker that has tripped. Quality triggers include a failed output guardrail, an abstention, a low-confidence classification, or a user repeating themselves or expressing frustration. Each trigger maps to one fallback, and the fallback chain should be short; three hops of increasingly weak models mostly adds latency.

The trap is assuming a fallback behaves like the primary. A different model may follow the same prompt differently, format output differently, or refuse things the primary allowed. Run your eval set against every fallback path, not just the main one, and alert when fallback traffic rises, because a fallback that silently carries a third of traffic for a week is now your primary, untested.

Can model confidence scores be trusted?

Not as given. Token probabilities and self-reported confidence carry real signal, but they are often miscalibrated, especially after instruction tuning, and verbal ratings cluster high. Trust them only after measuring how they relate to correctness on your own task.

There are several things people call a model's confidence, and they behave differently. Token log probabilities are the model's own probability for each token it generated; some providers expose them, others do not, and reasoning models often do not expose them for the hidden reasoning. Verbalised confidence is the model writing "I am 90% sure". Consistency is how often several sampled answers to the same question agree. External scores come from a separate verifier, a judge model or a retrieval relevance score.

Research gives a mixed picture. Work such as the 2022 paper Language Models (Mostly) Know What They Know found that large pretrained models can be reasonably calibrated on multiple-choice questions when you read probabilities over answer options. But instruction tuning and preference training are reported to make models more confident than they should be, and token probabilities on free text mostly measure how fluent a phrase is, not whether the fact is true. A wrong date can be generated with very high token probability because the format is predictable.

Verbalised confidence is the least reliable. Models tend to report high numbers, often between 80 and 95, regardless of difficulty, and the numbers shift with prompt wording. They are not measured probabilities; they are text that looks like one. Asking for confidence is still useful as a ranking signal if it correlates with correctness, which you can test.

In practice the most robust signals are usually not the model's own number. Self-consistency (sample five answers, measure agreement) costs more but often ranks errors well. Evidence-based signals (was the fact found in a retrieved source, did the citation check pass) are more trustworthy for factual tasks. Whichever you use, the question is not "is the score true" but "does a lower score mean more errors, and by how much", which is what calibration measures.

What is confidence calibration?

Calibration is the match between stated confidence and observed accuracy. A calibrated system is right about 80% of the time on answers it scores at 0.8. You measure it by bucketing predictions by score and comparing each bucket's accuracy with its average score.

Calibration answers a precise question: when the system says 0.8, how often is it right? If the answer is 0.8, the score is calibrated and you can use it directly in decisions, for example to estimate how many errors auto-approval will let through. If the answer is 0.55, the score is overconfident and any threshold you set from intuition will be wrong.

The standard measurement is a reliability diagram. Collect a labelled set, typically a few hundred to a few thousand examples, record the score and whether the answer was correct, group the scores into bins (0.0 to 0.1, 0.1 to 0.2 and so on), and plot average score against actual accuracy for each bin. A calibrated system sits on the diagonal. The expected calibration error (ECE) summarises the gap as a weighted average of the per-bin differences. Report it with the bin counts, because a bin with eight examples tells you very little.

Calibration is different from discrimination. Discrimination is whether higher scores go with more correct answers at all, often measured with AUROC. A score can rank errors perfectly and still be badly calibrated (every score is 0.95, but the low ones are the wrong ones), and that is fixable. A score that is calibrated on average but does not rank errors is useless for deciding when to abstain. Check discrimination first, then fix calibration.

Fixing calibration is a post-processing step. Temperature scaling and Platt scaling fit one or two parameters that reshape scores; isotonic regression fits a monotonic mapping from raw score to observed accuracy, which suits the lumpy scores LLM signals produce. Fit on one labelled split, check on another. Calibration is specific to the task, the prompt and the model version, so it must be redone when any of them changes, and monitored in production on a sample with human labels, because the input mix drifts.

How do you validate actions?

Validate a proposed action in layers, all in code: the arguments match the schema, the acting user is authenticated and authorised for this resource, business rules and limits hold, the target is in the state the model assumed, and the expected impact is within bounds.

A tool call proposed by a model should be treated like a request from an untrusted API client. The model may have misread the user, been misled by injected text, used stale data, or simply made up an ID. Validation is the code path that turns "the model wants to do X" into "X is allowed and still makes sense", and it runs every time, regardless of how confident the model sounded.

Work through the layers in order of cost, cheapest first. Schema: types, required fields, enums, formats and ranges, ideally enforced with constrained decoding and then checked again. Identity and authorisation: the action is evaluated against the end user's permissions on this specific resource, not the agent's service account; this is where most real damage is prevented. Business rules: limits, allowed time windows, account status, policy. Intent consistency: does the action match what the user asked for in this session, for example the amount and recipient the user actually typed, not ones that appeared in a retrieved document?

Then check state. Models act on what they last saw, which may be minutes old. Before updating a record, confirm it is still in the expected state, for example with a version number or an if-match precondition, so the agent does not overwrite a change a human just made. For destructive or large operations, compute the impact first: a dry run that returns how many rows would change, how many recipients would get the email, or the diff that would be applied. Reject or escalate when impact exceeds a bound.

Finally, make the call safe to retry with an idempotency key, and log the full decision: the proposed call, every check, the result and the reason. Return failures to the model as short structured reasons so it can explain or correct, rather than guessing a workaround.

Which actions need confirmation?

Ask a human to confirm actions that are irreversible, costly, externally visible, sensitive, or broad in scope. Show the exact resolved action, not the model's summary, and keep confirmations rare enough that people still read them.

Confirmation is a control with a budget. Every prompt to confirm costs the user time and attention, and if you ask too often they stop reading and click yes by reflex, which is called confirmation fatigue. The goal is to spend confirmations on the few actions where a mistake is expensive, and let everything else run with logging and the ability to undo.

A useful test asks five questions about an action. Is it irreversible or hard to undo (deleting data, sending a message, executing a payment)? Is it costly above a threshold? Is it external, meaning someone outside the organisation sees it (an email to a customer, a public post, a pull request to an open repository)? Does it touch sensitive data or permissions (sharing a document, changing access, exporting records)? Is it broad, affecting many records or recipients at once? One strong yes is usually enough to require confirmation; reversible, internal, single-record reads and drafts usually need none.

How you confirm matters as much as when. Show the resolved action: the actual recipient list, the amount with currency, the diff, the number of affected rows, generated from the tool arguments by code. Do not show the model's paraphrase, because the paraphrase and the arguments can disagree, and an injected instruction can make the paraphrase misleading. For high-stakes actions, require the user to type or select a key value rather than press one button. In voice interfaces, read back the critical values and require an explicit yes.

Confirmation is not a substitute for authorisation. A user confirming an action they are not allowed to take should still be refused. And for actions so dangerous that a busy user might confirm them carelessly, use a second approver or remove the capability from the agent entirely.

How do you prevent catastrophic mistakes?

Assume the model will eventually propose the worst action it can, and make that action impossible or small: least-privilege tools, hard caps and rate limits, approvals for irreversible steps, dry runs, reversible operations, monitoring with automatic stops, and a tested kill switch.

Catastrophic mistakes rarely come from a single bad answer. They come from an agent with broad permissions taking a plausible but wrong action at scale: deleting a production table instead of a test one, emailing a draft to every customer, issuing thousands of refunds in a loop. The prevention strategy is to bound the blast radius, the maximum damage any single action or run can do, so that the worst case is survivable regardless of what the model decides.

Start by removing capability. Least privilege means the agent's credentials can only touch what its job needs: read-only database users for analysis agents, scoped API tokens, a sandbox for code execution, no generic shell or SQL tool in production paths. An action the agent cannot take is a mistake it cannot make. Then add limits: per-action caps (refunds up to an amount), per-run and per-day budgets (total value, number of writes, number of messages), and rate limits that slow a runaway loop down enough for monitoring to catch it.

Next, make mistakes reversible. Prefer soft deletes, versioned writes and staged changes over destructive operations. Send emails through a short delay queue that can be cancelled. Apply bulk changes in batches, starting with a small canary batch. For changes that cannot be reversed, require a dry run that shows the impact and a human approval, and for the most dangerous, a second approver.

Finally, detect and stop. Monitor actions, not just errors: a spike in writes, refunds or messages per minute is an incident even if every call succeeded. Wire those monitors to automatic circuit breakers that pause the agent, and keep a manual kill switch that on-call staff have actually used in a drill. Keep an audit trail detailed enough to find and reverse every action from a bad run.

What is a policy engine?

A policy engine is a component that takes a proposed action and its context (who, what, which resource, when, how much) and returns allow, deny or escalate according to rules written as code or data, separate from the agent and the model.

As an agent gains tools, the rules about what it may do multiply: refunds below a limit, no external email after hours, only managers can change prices, nothing touches accounts flagged for fraud. Scattering those rules across tool implementations and prompts makes them hard to find, test and change. A policy engine pulls them into one place with one interface: given a request, return a decision and a reason.

The request usually carries four kinds of input. The principal: the end user, their role and attributes, and the agent acting for them. The action: the tool and its arguments. The resource: what is being acted on and its attributes, such as owner, sensitivity or status. The context: time, channel, risk signals, cumulative totals for the session or day. Policies are rules over those inputs. A good engine evaluates them deterministically, defaults to deny when no rule allows, and returns a machine-readable reason.

You can build this as plain functions in your own language, which is fine for a handful of rules. As rules grow, general-purpose policy engines help: open-source options such as Open Policy Agent (with its Rego language) or Cedar let you write policies as data, version them, test them in isolation and evaluate them in milliseconds. Whatever you choose, the important properties are that policies are reviewed like code, have unit tests with explicit allow and deny cases, and log every decision for audit.

Keep the model out of the decision. It is tempting to let a model judge whether an action "seems appropriate", and a model check can be a useful extra signal for escalation. But the policy engine itself should be deterministic, so the same request always gets the same answer and an injected prompt cannot argue with it. Treat a policy change as a release: test it against recorded past requests to see which decisions would flip before you deploy it.

What is a kill switch, and how do circuit breakers keep an AI system safe?

A kill switch is a fast, pre-built way to stop an AI feature or agent, fully or partly, without a deploy. A circuit breaker is the automatic version: it trips when error rates or action volumes cross a threshold and routes traffic to a safe fallback.

AI incidents often look different from ordinary outages. Every request succeeds with a 200 status, but the model has started giving harmful answers after a provider update, or an agent is following an injected instruction across many conversations, or a cost spike shows a loop running in thousands of sessions. Normal health checks stay green. You need a way to stop the behaviour within minutes, and you need to have built it before the incident.

A kill switch is a runtime flag, read on every request, that disables a capability. Good kill switches are granular: turn off one tool ("refunds"), one agent, one model route, one customer tenant, or the whole AI feature, each falling back to a defined safe behaviour such as a static message, a non-AI path or a human queue. They should not require a deploy, should take effect in seconds, and should be operable by on-call staff without special access. Test them in a drill; a switch that has never been flipped tends to have a hidden dependency on the system it is meant to disable.

A circuit breaker, borrowed from distributed systems, trips automatically. It watches a signal over a sliding window and, when the signal crosses a threshold, it opens: requests stop going to the risky path and take the fallback instead. After a cool-down it lets a few trial requests through (half-open) and closes again if they succeed. For AI systems, watch more than errors. Useful signals include guardrail block rate, abstention rate, action count or value per minute, tokens or spend per session, and user complaints or thumbs-down rate.

The trade-off is false trips against slow response. Thresholds that are too tight shut off a working feature during a normal traffic surge; thresholds that are too loose let an incident run. Base thresholds on observed baselines with a margin, alert a human whenever a breaker opens, and make breakers that guard irreversible actions fail toward pausing rather than continuing.

How do you know your guardrails actually work?

Measure each guardrail like a classifier: how often it catches the failures it targets (recall), how often it blocks legitimate traffic (false positive rate), and what latency and cost it adds. Test it with labelled benign and adversarial sets, then keep measuring in production.

A guardrail that has never been measured is a belief, not a control. Teams often add a filter, see that it blocks the obvious test case, and move on. Months later they discover it blocked 4% of legitimate customer questions, or that a reworded attack passes straight through. Every guardrail is a classifier making a decision, and classifiers have error rates you can measure.

Build two test sets per guardrail. A benign set of realistic legitimate inputs, ideally sampled from production, including the awkward edge cases near the boundary (a nurse asking about overdose thresholds, a security engineer asking how an exploit works). A harmful set of the things the guardrail should catch, including variations from red teaming: paraphrases, other languages, encodings, instructions split across turns or hidden in documents. From those, compute recall on the harmful set and the false positive rate on the benign set. Report both, since either alone can be made perfect by blocking everything or nothing.

Then measure the cost of running it. Model-based guardrails add latency, often more than the rest of the request on simple queries, and they add per-request cost. Record p50 and p95 latency per guardrail. If a guardrail runs in series before the main call, its latency is paid on every request; running cheap checks in parallel with generation, and blocking only at the end, can hide much of it.

In production, log every guardrail decision with enough context to review. Sample blocked and passed requests for human review each week to estimate real-world error rates, and track the block rate over time; a sudden rise usually means either an attack or a model or product change that broke the guardrail's assumptions. Add every confirmed miss to the harmful set so the next version is tested against it.