The GenAI Field Guide

Questions that test real understanding

The deeper trade-offs that distinguish a demo from a dependable system.

This chapter is about the questions that separate a GenAI demo from a system people can depend on. Each one starts from something that sounds true and turns out to be only partly true: a better prompt is better, more context helps, bigger models win, more agents can do more. The answers explain why these intuitions break down in production, and what experienced teams do instead.

The mental model running through every answer is that a language model is a probabilistic component inside a deterministic system. It is excellent at interpreting messy language and producing fluent text, and it is right most of the time, not all of the time. Reliability comes from the system around it: evals that measure how often it is right, tools and code that bound what a wrong output can do, durable state so failures do not repeat side effects, and traces that show where a bad answer came from.

The Basic questions cover the everyday surprises: prompt edits that regress other cases, context that distracts, small models that win, demos that mislead, multi-agent designs that cost more than they deliver, and tools as the real safety boundary. The Advanced questions cover the engineering that makes a system dependable: why error rates compound over long tasks, which rules must live in code, state and idempotency, observability, what production-grade means, how to stay up when the model fails, where to draw the line between model and code, and why quality drifts after launch even when nobody touched the code.

Why can a better prompt make a system worse?

A prompt edit that fixes the case you were looking at changes behaviour on every other input too, so without a broad eval it can silently break cases that used to pass, raise cost, or bury a rule that mattered.

A prompt is not a local patch. Every instruction is read on every request, so adding a sentence to fix one complaint changes the distribution of outputs for all inputs. The model has no notion of "this rule only applies to the refund case" unless you say so precisely, and even then the new text shifts emphasis away from what was already there. A prompt that looks better when you read it, or that fixes the three examples on your screen, has not been shown to be better at all.

Several mechanisms cause the damage. Dilution: a longer prompt gives each rule a smaller share of attention, so a critical constraint that used to be followed now slips occasionally. Conflict: a new instruction ("always be thorough") contradicts an old one ("answer in two sentences"), and the model resolves the tension differently from request to request. Over-correction: a strongly worded fix ("NEVER mention competitors") makes the model refuse legitimate questions that merely touch the topic. Format drift: a new example changes the output shape that a downstream parser depends on.

There are also costs that never show up in a side-by-side read. More instructions mean more input tokens on every call, and if the edit lands before the stable part of the prompt it can break prompt caching, raising both cost and latency. A prompt tuned on one model version can also behave differently on the next, so the "improvement" may be specific to a model you are about to replace.

The fix is to treat prompts like code with tests. Keep the prompt in version control, run every edit against a fixed eval set that covers the main task types and the known hard cases, and compare pass rates per slice, not just the overall average. An edit is better only if it improves the target slice without dropping others beyond noise.

Why isn't more context always better?

Every extra token competes for the model's attention, adds cost and latency, and can introduce outdated or conflicting material; past a point, adding context lowers accuracy rather than raising it.

A large context window means the model can accept a lot of text, not that it uses all of it well. Research such as the 2023 paper Lost in the Middle showed that models tend to use information near the start and end of a long input more reliably than information buried in the middle. Later work on long-context benchmarks has found that accuracy often falls as input length grows, even when the needed fact is present. The effect varies by model and task, but the direction is consistent enough to plan around.

The bigger problem is usually what the extra text contains. Pasting in ten retrieved documents when two are relevant adds distractors: passages that look related and pull the answer towards the wrong policy, the wrong customer or the wrong product version. If the context holds both the 2024 and the 2026 version of a manual, the model has no reliable way to know which wins unless you tell it. Contradictions in context are a common root cause of answers that are confidently wrong.

There is also a plain economic cost. Input tokens are billed on every call, and the time to process the prompt grows with its length, which pushes up time to first token. A chat that resends a long history and twenty documents on every turn can cost several times more than one that sends a summary and the three relevant passages, with no gain in quality.

The goal of context engineering is the smallest set of high-signal material that lets the model do the task: the instructions, the facts it needs, and nothing it would have to ignore. More context is worth it when the task genuinely needs it, such as reasoning across a whole contract, and when the material is clean and current.

Why can a small model beat a large model in a system?

In a well-built system the model is one part among retrieval, tools, constraints and checks, so a small model with the right context and a narrow task often matches or beats a large model working with poor context.

Model benchmarks measure what a model can do from its own knowledge and reasoning on general tasks. Most production tasks are narrower: classify this ticket, extract these six fields, answer from this policy. For those, the deciding factor is usually whether the model sees the right information and is asked a well-defined question, not how much general capability it has in reserve.

A large model asked a vague question with stale context will produce a fluent, plausible answer that is wrong. A small model handed the exact policy clause, a strict output schema and one clear job will usually get it right, and a validator can catch the cases where it does not. Retrieval supplies the facts, tools do exact work such as arithmetic and lookups, and structured output removes formatting as a source of error. Each of those shrinks the gap between small and large.

Small models also bring system-level advantages. They are faster, which matters in voice and interactive use, and cheaper, which lets you afford things that improve quality: running the task twice and comparing, adding a verification step, or covering more traffic with evals. A small model can also be fine-tuned or distilled for one narrow task, where it often reaches the quality of a much larger general model.

This does not mean small always wins. Open-ended reasoning, long multi-step planning, ambiguous user requests and tasks with little supporting context still favour larger models. The practical answer is to measure on your own eval set and often to combine them: route easy cases to a small model and escalate hard or low-confidence cases to a larger one.

Why are evals more important than demos?

A demo shows that a system can succeed on cases someone chose; an eval measures how often it succeeds on representative cases, including the hard and rare ones, and whether a change made it better or worse.

A demo is a sample of one, picked by someone who wants it to work. It proves the system can produce a good answer. It says nothing about the rate of bad answers, which is what users, support teams and compliance reviewers actually experience. A system that is right 80% of the time can give a flawless demo every time if the presenter knows which questions to ask.

An eval is a fixed set of realistic cases with a way to score each output. Because the set is fixed, you can run it before and after any change and see the difference. Because it is representative, the score approximates what users will see. Because it includes slices for rare and hard inputs (scanned documents, mixed languages, ambiguous requests, adversarial text), it surfaces failures that a demo would never touch.

Evals matter most for change. Models are updated by providers, prompts are edited, retrieval indexes are rebuilt, and each of these can quietly move quality. Without a regression suite, the first signal is a user complaint weeks later. With one, a drop on the scanned-invoice slice from 91% to 78% shows up in CI before the change ships.

Demos still have a role: they communicate what the product does and help stakeholders form expectations. The mistake is letting a demo stand in for evidence. A useful habit is to show the eval numbers alongside every demo, including the failure rate and an example of a failure.

Why can multi-agent systems underperform?

Splitting work across agents adds handoffs, each of which can lose context, duplicate effort or introduce disagreement, so unless the subtasks are genuinely independent the extra coordination costs more than it gains.

A multi-agent system breaks a task across several model instances with different roles, such as a planner, researchers and a reviewer. The appeal is the same as dividing work among people. The catch is that agents communicate only through text they pass to each other, and every handoff is a lossy summary. A sub-agent rarely sees all the context the orchestrator had, so it can make choices that are locally sensible and globally wrong.

Several failure modes recur. Context loss: a constraint stated to the planner never reaches the agent doing the work. Duplicated work: two agents search for the same thing, doubling cost. Conflicting outputs: parallel agents produce incompatible pieces, such as two code changes that each assume the other file is unchanged. Debate loops: a reviewer and a writer pass a draft back and forth without converging. Compounding: more steps means more chances for any single step to go wrong.

There is also a cost multiplier. Each agent re-reads instructions and context, so token use can be several times that of a single agent on the same task. Latency grows with every sequential handoff. Debugging is harder because a bad result may come from any agent or from the message between them.

Multi-agent designs do pay off in specific shapes: breadth-first research where subtopics are independent and can run in parallel, tasks that need more total context than one window holds, or isolation for security reasons, such as a sub-agent that reads untrusted content without access to dangerous tools. Outside those shapes, a single agent with good tools and a clear plan usually wins.

Why is tool design so important?

Tools define the full set of actions an agent can take, so their names, parameters, limits and error messages decide both how often the model uses them correctly and how much damage a wrong call can do.

When a model calls a tool, it chooses from the menu you gave it, using only the names, descriptions and parameter schemas it can see. It cannot read your code or your intentions. If two tools have overlapping descriptions, it will pick the wrong one some of the time. If a parameter is a free-form string where an enum would do, it will invent values. Tool definitions are effectively part of the prompt, and often the most important part.

Tool design is also the main safety boundary. A model that can only call issue_refund(order_id, amount) with a server-side cap cannot transfer arbitrary money, however it is manipulated. A model with a generic run_sql(query) or http_request(url, body) can do anything the credentials allow. Narrow, purpose-built tools encode business rules in code; broad tools leave those rules to the model's judgement, which prompt injection can subvert.

Good tools share a few traits. They match a task, not an API endpoint: find_customer_by_email beats exposing a raw search with fifteen filters. They validate inputs and enforce limits server-side. They return compact, relevant results rather than whole database rows. Their errors explain what to do next ("order 1182 is already refunded; no action taken") so the model can recover instead of retrying blindly. Write tools are idempotent where possible, so a retry does not repeat a side effect.

The trade-off is coverage against control. Many narrow tools are safer and easier to use correctly but take effort to build and can bloat the tool list, which itself confuses the model. Group related actions sensibly, keep the visible tool set small for each task, and expose broad capabilities only inside a sandbox.

Why do agent error rates compound?

When a task needs many steps and each must succeed, the end-to-end success rate is roughly the product of per-step rates, so 95% per step over ten dependent steps gives only about 60% overall.

Consider an agent that must find the right file, read it, identify the bug, edit the code, run the tests and report. If any step goes wrong and nothing catches it, the task fails. If each step independently succeeds with probability p, a chain of n steps succeeds with probability p to the power n. At 95% per step, ten steps give 0.95^10, about 60%. Twenty steps give about 36%. At 99% per step, twenty steps still give about 82%. Per-step accuracy that sounds excellent becomes mediocre over a long task.

The arithmetic is a simplification, and the real picture is both worse and better. Worse, because errors propagate: a wrong assumption in step two shapes every later step, and the model tends to stay consistent with its own earlier output rather than question it. A mistaken file path does not just fail one step; it sends the agent down a wrong branch that wastes the rest of the budget. Steps are also not equally hard, and the weakest step dominates.

Better, because not every error is fatal. An agent that runs tests can notice a failed edit and try again. A validation step can reject bad output before it is used. If a step that fails 5% of the time is followed by a check that catches 80% of those failures and triggers a retry, the effective failure rate for that step drops to about 1%. This is why the reliability of long-horizon agents depends so much on feedback signals: tests, schema validators, tool errors and explicit verification.

The design consequences follow directly. Shorten chains by doing deterministic work in code instead of asking the model. Add checks after the steps most likely to fail. Make each step's success observable so the agent and your traces can see when something went wrong. Measure end-to-end success on tasks of realistic length, because per-step accuracy alone will flatter the system.

Why keep critical actions deterministic?

Rules that must always hold, such as permissions, limits, money calculations and irreversible actions, belong in code because code enforces them exactly every time, while a model follows them only most of the time and can be talked out of them.

A language model is a probabilistic component. Given the same instruction, it follows it with high probability, not certainty, and that probability shifts with phrasing, context length, model version and adversarial input. For a summary, 98% adherence is fine. For "never refund more than the order total" or "only show records this user owns", a 2% violation rate is a financial or legal incident at scale. At 100,000 requests a day, 2% is 2,000 violations.

Prompt injection makes this sharper. Any text the model reads, whether a customer email, a retrieved web page or a tool result, can contain instructions. A model that enforces policy through its own judgement can be persuaded to skip that policy. Code that checks the user's permission, the refund cap and the approval status cannot be persuaded, because it does not read the email.

The pattern is model proposes, code disposes. The model interprets the request and suggests an action with parameters. Deterministic code then validates the proposal against hard rules: is the caller allowed, is the amount within limits, does the target exist, is this action reversible, does it need human approval. Only then does it execute. The model never holds the credentials or calls the payment API directly.

This also helps with audit and debugging. When a rule is in code, you can point to the line that enforced it, write a unit test for it, and prove to an auditor it holds. When a rule lives in a prompt, the only evidence it held is the outputs you happened to check. The trade-off is engineering effort and some loss of flexibility: code rules must be written for each action, and edge cases the code did not anticipate will be rejected and need a human path rather than a model improvising.

Why is state management critical?

Agent runs are long, pause for approvals and fail midway, so without durable state and idempotent actions a resumed or retried run either loses progress or repeats side effects like sending an email twice.

A chat reply that fails can simply be retried. An agent run is different: it may take minutes or hours, call a dozen tools, wait a day for a human approval, and survive a deploy or a crashed worker in the middle. If the only record of progress is the model's context window in a process's memory, any of those events loses the run. Restarting from scratch wastes cost and, worse, re-executes actions that already happened.

Two things are needed. Durable state: after each meaningful step, persist what the run has done, its intermediate results and what it is waiting for, in a database rather than in memory. This is checkpointing. On resume, the harness loads the checkpoint and continues from the next step instead of replaying the conversation and hoping the model makes the same choices. Idempotent side effects: every action that changes the outside world carries an idempotency key, a stable identifier for this specific intended action, so that if it is attempted twice, the second attempt is recognised and skipped.

The classic bug is the window between doing something and recording that it was done. A worker sends an email, then crashes before saving "email sent". On retry it sends again. The fix is to record the intent with a key before acting, pass that key to the downstream system where it supports one, and mark completion after. Many payment and messaging APIs accept idempotency keys for exactly this reason. Where the downstream system does not, you can only narrow the window, so prefer a provider that does for anything costly to repeat.

State also has to be explicit about decisions. If a run resumes after an approval, the approval decision, who made it and when belong in state, not just in a chat message. Where state lives depends on scale: a relational database is the default, with a queue or durable workflow engine for long-running orchestration. The trade-off is design effort up front, but adding durability after a duplicate-payment incident is much harder.

Why is observability essential?

A bad answer can come from the prompt, retrieval, a tool, the model or the data, and only a trace of every step with its inputs, outputs, tokens and latency lets you find which one, and how often it happens.

In traditional software, an error usually throws an exception at the line that broke. In an LLM system, most failures are silent: the request returns 200 with a fluent, wrong answer. The cause could be any of a dozen components. Retrieval may have returned a stale document. A tool may have timed out and the model guessed instead. The prompt template may have dropped a variable. The provider may have changed model behaviour. Without a record of what happened inside the request, you are left re-running it and hoping to reproduce the problem, which non-determinism often prevents.

Observability for LLM systems means capturing a trace per request: a tree of spans, one per step, each with its inputs, outputs, timing, token counts, model and prompt version, and errors. For a retrieval step, that means the query and the returned document ids and scores. For a tool call, the arguments and the result. For a model call, the rendered prompt or a reference to it, plus the output. With that, a bad answer becomes a debugging session of minutes rather than days.

Traces also feed everything else. Aggregated, they give cost per feature, latency percentiles, tool error rates and retrieval hit rates. Sampled and scored, they become online evals. Failed traces, labelled, become regression cases. They also form part of the audit trail for actions taken on a user's behalf.

The trade-offs are privacy and volume. Prompts and outputs often contain personal data, so traces need access controls, redaction and retention limits that match the data's sensitivity. Full payload logging at high traffic is expensive, so teams commonly keep metadata for every request and full payloads for a sample plus every error. Open standards such as OpenTelemetry now include semantic conventions for generative AI spans, which reduces lock-in to one tracing vendor, though those conventions are still evolving.

What belongs in a production-grade GenAI system?

Beyond the model and prompt: measured quality with evals, controlled actions enforced in code, scoped data access, tracing, cost and rate limits, failure handling with fallbacks, and a named owner who responds when it breaks.

A prototype answers the question "can a model do this?". A production system answers "will it keep doing this acceptably, for real users, with real data, when things go wrong, and will we know when it stops?". The model call is usually a small fraction of the code. Most of the work is in the surrounding system that makes the model's behaviour measurable, bounded and recoverable.

The components fall into a few groups. Quality: an eval set covering real traffic slices, a regression gate in CI, and online sampling of production outputs. Control: tools scoped to tasks, permission and limit checks in code, human approval for high-impact actions, and an audit log of every action taken. Data: access scoped to the calling user or tenant, retrieval that respects permissions, redaction where needed, and retention rules. Operations: tracing, dashboards for quality, cost and latency, rate limits and per-user budgets, timeouts and retries, provider fallback, and alerting. Ownership: someone on call, a runbook, and a process for turning incidents into eval cases.

Not every system needs all of it at full strength. An internal summariser with no write tools needs evals, tracing and cost limits but not an approval workflow. A customer-facing agent that can move money needs every item. The useful exercise is to go through the list for your system and write down, for each item, either how it is handled or why it is not needed.

The common gap is not technology but ownership. Many systems launch with good engineering and then degrade quietly because nobody is responsible for reading the dashboards, reviewing failed traces or updating the eval set when the product changes. Production-grade means someone is accountable for quality after launch, with time allocated to it.

How can a system stay reliable when a model fails?

Design for failure as a normal case: verify important outputs, bound what any single output can do, retry and fall back for outages, and give the model and the system a safe way to abstain and hand off to a person.

Models fail in two very different ways, and they need different defences. Operational failures are visible: timeouts, rate-limit errors, provider outages, truncated or malformed output. Quality failures are invisible: a well-formed, confident answer that is wrong. Reliability engineering for GenAI has to cover both, and the second is harder because nothing throws an error.

For operational failures, the toolkit is familiar from distributed systems. Set timeouts on every call. Retry transient errors with exponential backoff and jitter, but cap attempts. Validate structured output against its schema and retry once with the validation error if it fails. Keep a fallback path, such as a secondary model or provider, a cached answer, or a simpler non-AI flow. Use circuit breakers so a failing provider is skipped rather than hammered.

For quality failures, the principle is to verify what matters and bound the rest. Check outputs that drive decisions: does the cited passage actually support the claim, do extracted numbers reconcile, is the proposed action allowed. Where you cannot verify, limit the impact: route the output to a draft rather than sending it, require approval, or cap the amount. Give the model an explicit way to say it does not know, and treat that as a valid outcome rather than a failure to be prompted away.

Abstention and handoff are what make this work for users. A system that sometimes says "I can't confirm this; I've sent it to a specialist" is more trustworthy than one that always answers. The trade-off is coverage: every abstention is a case a human must handle. Tune the thresholds on an eval set so the abstention rate is affordable and the error rate on answered cases is acceptable, and watch both over time.

Which parts of a system should be probabilistic?

Use models where the task needs interpretation of messy language or generation of new text; use exact code for permissions, calculations, state changes and anything with a single correct answer that a program can compute.

A useful way to split a system is to ask, for each step, whether there is a single correct answer that a program could compute. If there is, such as a tax amount, a date difference, whether a user owns a record, or whether a field matches a pattern, code does it better: exactly, instantly, cheaply and testably. If there is not, such as understanding what a frustrated customer is asking, summarising a long thread, or drafting a reply in the right tone, a model is the right tool because no practical program could do it.

Most real tasks mix both, and the craft is in drawing the boundary. A model reads an invoice and extracts the line items (interpretation). Code sums them and checks the total (calculation). A model classifies a ticket's intent (interpretation). Code routes it to a queue based on the label and the customer's plan (rules). A model drafts the refund explanation (generation). Code decides whether the refund is allowed and executes it (permission and transaction).

The boundary is usually a structured interface. The model's output is a typed object, validated against a schema, that code then acts on. This makes the probabilistic part easy to evaluate in isolation and keeps its errors from flowing directly into actions. It also makes upgrades safer: you can change the model without touching the rules.

There are grey areas. Fuzzy matching of names, ranking search results and detecting anomalies have learned components that are probabilistic but do not need a language model; a classical model or a heuristic may be cheaper and more predictable than an LLM. Conversely, some rule systems are so large and full of exceptions that a model with the rules retrieved does better than brittle code, provided its output is checked. Decide by the cost of an error and by whether the result can be verified.

Why can quality drift after launch when nothing in your code changed?

Your code is only one input: the provider's model, your indexed data, user behaviour and upstream tools all change over time, so a system that passed its evals at launch can degrade without a single commit.

A GenAI system depends on several things you do not version in your repository. The model can change: hosted providers release new versions, and aliases that point to "the latest" version move to them. Pinned versions are eventually deprecated, forcing a migration, and providers do not always document every change to serving behaviour. The data changes: documents are added, edited and left stale in the retrieval index, and a broken re-index job can serve months-old content. Users change: a new feature, a marketing campaign or a new customer segment brings questions the eval set never covered. Tools and APIs change: an upstream service renames a field or slows down, and the agent starts guessing.

None of these shows up as a code diff or a failed unit test. The usual first signal is a rise in complaints, thumbs-down ratings or escalations weeks later, by which time the cause is hard to pin down. This is why observability and evals have to continue after launch, not just gate the release.

Three habits catch drift early. Pin and record versions: use explicit model versions where the provider offers them, and record model, prompt and index version on every trace so a change in quality can be lined up with a change in inputs. Re-run evals on a schedule, not only on code changes, so a provider update or data change is caught even when no one deployed. Monitor production slices: sample real traffic, score it with the same graders as offline, and watch the distribution of incoming requests for new topics or languages that the eval set does not cover.

The trade-off with pinning is that you also miss improvements and eventually face a forced migration. Treat a model upgrade like any other change: run the full eval set on the new version, compare per slice, and roll it out gradually.