The GenAI Field Guide

Practical architecture decisions

Turn vocabulary into concrete choices for your own product.

This chapter turns the vocabulary of the rest of the guide into decisions. Most GenAI architecture questions have the same shape: a simple option that works for most teams, a more complex option that is right under specific, measurable conditions, and an expensive mistake of adopting the complex one too early. Model routers, multi-agent systems, dedicated vector databases and long-term memory all solve real problems, and every one of them adds a component that must be evaluated, observed, secured and paid for.

The mental model is that every component must earn its place with a measured benefit. Default to the fewest moving parts: one model, a deterministic workflow with model steps, whole documents when they fit, your existing database for state and vectors, no extracted memory. Then let evals, traces and cost data show where complexity pays. Separate decisions by how hard they are to reverse. Model choice and prompts are cheap to change, so make them quickly and put them behind configuration. Where state lives, what the system may change, how retries stay safe and where people approve actions are costly to change later, so decide them deliberately and early.

The questions build in that order. The Basic questions cover structural choices: one model or several, retrieval or whole documents, workflow or agent, one agent or many, memory, vector storage, conversation state, and synchronous versus background execution. The Advanced questions cover operating the system: where human approval goes, how retries work, how to survive provider outages, how to debug a bad answer, what to verify before launch, how to design a new system from a blank page, and how to keep it adaptable as models change.

One powerful model or several models?

Start with one capable model that meets your quality bar on your own eval set. Add a cheaper or faster model only for a task slice where evals show it matches quality and the saving justifies the extra routing, testing and monitoring.

Using one model everywhere has benefits that do not show up on a pricing page: one prompt style to maintain, one set of quirks to learn, one provider contract and data policy, and one eval baseline. While you are still discovering what the product needs, these matter more than per-token savings. Most early systems do not have enough traffic for model spend to outweigh the engineering time a multi-model setup consumes.

Several models become worth it when traffic is uneven in difficulty. In a support assistant, perhaps 70% of requests are classification, field extraction or short FAQ answers that a small model handles as well as a large one, while 30% need multi-step reasoning. Sending the easy slice to a smaller model cuts cost and latency for most of the traffic. Other legitimate reasons are a specialised capability (embeddings, speech, vision, a fine-tuned classifier), a residency rule that forces a particular deployment, or a latency budget the large model cannot meet.

The costs are real. Each model needs its prompts tuned, its own eval runs and its own regression checks whenever the provider updates it. Two patterns are common. A router picks a model up front from features of the request. A cascade tries the cheap model first and escalates when a check fails, such as schema validation or a verifier verdict. A router is itself a classifier that can be wrong: a misrouted hard question reaches a weak model and the user sees a poor answer with no error anywhere in the logs.

Make the split with data, not intuition. Log task types in production, build an eval slice per task type, run the candidate models on each slice, and compare quality, p95 latency and cost per task. Move only the slices where the smaller model is within your tolerance, keep the large model as the escalation path, and monitor misroutes by sampling routed traffic.

RAG or the whole document?

Send the whole document when the relevant material fits comfortably in the context window and the question may need any part of it. Retrieve when the corpus is too large, changes often, has per-user permissions, or each question needs only a small part.

Long context windows changed this decision. A model that accepts hundreds of thousands of tokens can read a whole contract, policy manual or code module in one request. Sending the whole document removes a class of failures: no chunking that separates a clause from its exception, no retrieval step that misses the paragraph that mattered, and the model can connect a definition on page 2 with an obligation on page 40.

Whole-document prompting has costs. You pay input tokens for the full document on every call, although prompt caching cuts this when the same document is reused across calls (how much varies by provider). Time to first token grows with input length. Quality on long inputs is uneven: models use material in the middle of a very long context less reliably, and irrelevant text competes for attention, an effect often called context rot. A large window does not mean every part of it is used equally well.

RAG (retrieval-augmented generation, fetching relevant passages and adding them to the prompt) wins when the corpus is larger than any window, when it changes daily, when different users may see different documents, or when answers must cite specific passages. It also keeps cost per question roughly flat as the corpus grows.

Many good systems combine the two. They retrieve to decide which documents matter, often with metadata filters, then send those documents whole rather than as small chunks. This keeps the cross-section reasoning of full context while scaling to large collections.

Decide by testing, not by window size. Build a set of questions with known answers, run both approaches, and compare accuracy, cost per question and p95 latency.

Agent or deterministic workflow?

Use a deterministic workflow when the steps are known in advance, and let a model choose the path only where the next step genuinely depends on what it discovers. Most dependable systems are workflows with a few bounded agent steps.

The difference is who controls the flow. In a workflow, your code fixes the sequence and model calls are steps inside it: extract fields, classify, draft a reply. In an agent, the model decides in a loop which tool to call next and when the task is finished. Both use the same models; the question is whether the path is written in code or chosen at run time.

Workflows are predictable. You can test each step on its own, bound the worst-case cost and latency, reason about what happens when a step fails, and show an auditor exactly what ran. Agents are flexible, but every step is a probabilistic decision and errors compound: if each decision is right 95% of the time, ten decisions in a row all go right only about 60% of the time. Agent cost and latency vary per run, and evaluating them means judging whole trajectories, not single outputs.

Agents earn their place when the path cannot be enumerated: investigating why a reconciliation failed, debugging code, researching across sources you cannot list in advance, or any task where the number of steps depends on what is found. Coding agents are the clearest case.

The strongest pattern is usually a hybrid. Keep the main path as code and make only the open-ended exception an agent step, with read-only tools, a step limit and a structured result that code then acts on. A simple test helps: if you can draw the flowchart and it has a manageable number of branches, write it as code.

One agent or multiple agents?

Start with one agent with good tools and well-managed context. Split into several only when subtasks are genuinely independent and can run in parallel, or one context cannot hold the work, and confirm on evals that the split improves quality or time.

In a multi-agent design an orchestrator delegates subtasks to subagents, each with its own context, prompt and tools, and combines what they return. The appeal is specialisation and parallelism. The cost is that every handoff loses information: a subagent sees only the brief the orchestrator wrote, and the orchestrator sees only the summary that comes back.

That information loss causes the typical failures. Subagents duplicate work, make conflicting assumptions (two agents writing parts of one program with different designs), or miss constraints that were obvious in the parent context. Token use multiplies because each agent re-reads its own context, and reported multi-agent research systems use several times the tokens of a single agent. Debugging spans several traces that must be stitched together.

Multi-agent works best on breadth-first, read-heavy, parallel tasks: searching five independent data sources, reviewing twenty files separately, or comparing vendors where each comparison stands alone. Each subagent explores widely and returns a compact finding. A separate context also acts as isolation: a subagent can read a large log file and hand back a short summary, keeping the main context clean.

It works poorly on tightly coupled tasks where each decision depends on the others, such as writing one coherent document or one module of code. Designs that copy an org chart (a planner, a coder, a tester and a manager talking to each other) usually add handoffs without adding independent work.

Treat the split as an experiment: run the same eval with one agent and with several, and compare success rate, tokens per task and wall-clock time.

Should my app have memory?

Only when facts kept from earlier sessions measurably improve later work, and only with user visibility and control. Most apps need durable conversation history and explicit settings, not an open-ended store of whatever the model decides to remember.

Three things get called memory. Conversation history is what was said in the current thread. Explicit user data is settings and profile fields the user set on purpose and your database stores. Long-term memory is facts the system extracts from conversations and injects into future ones. The first two are ordinary engineering. The third is the real decision.

The value is real for some products. Remembering that a user wants weekly reports as a table, works in a particular codebase, or writes in British English saves them repeating it. Assistants used daily for ongoing work benefit most; one-off question answering benefits little.

The risks are also real. A wrong or stale memory persists and quietly biases every later answer. Extracted facts may be misreadings of what the user meant. Memory poisoning occurs when injected text in a document or tool result causes the system to store an instruction that affects future sessions. Privacy obligations follow the data: anything stored must be covered by retention rules, data-subject deletion requests and tenant isolation, and sensitive categories such as health details or one-time secrets should never be retained.

If you add memory, store small, typed, attributable records: what was learned, from which conversation, when, and when it expires. Write memories only from the user's own messages, not from retrieved documents or tool output. Scope them to the user and tenant, show them in the interface with edit and delete, retrieve only those relevant to the current task, and run your evals with memory on and off to confirm it helps.

Do I need a vector database?

Often not at first. Keyword search, or vectors stored in a database you already run such as PostgreSQL with a vector extension, handles many corpora well. Adopt a dedicated vector database when measured scale, throughput or feature needs outgrow that.

A vector database stores embeddings (lists of numbers representing meaning) and finds the nearest ones to a query quickly using approximate indexes such as HNSW, a graph-based index. Dedicated products add managed scaling, sparse and hybrid search, multi-tenancy features and storage tiers.

Many teams do not need one on day one. Keyword search with BM25 ranking is strong exactly where embeddings are weak: product codes, error messages, names and rare terms. For thousands to a few million chunks, a vector extension in PostgreSQL (pgvector is the common one) provides approximate search, and you keep transactions, joins, backups and row-level security in one place. Document metadata and access rules are already in that database, so filtering by tenant or permission is a WHERE clause rather than a sync job to a second system.

A dedicated store pays off with hundreds of millions of vectors, high query throughput that would load your transactional database, or features your current store lacks. Its cost is another system to keep in sync. Stale or undeleted vectors are a correctness problem and, when a user's data should have been erased, a privacy problem.

The store is rarely what limits retrieval quality. Chunking, the embedding model, hybrid keyword-plus-vector search and reranking decide most of it. Measure recall@k (how often the right passage is in the top k results) on labelled questions before changing infrastructure.

Where should conversation state live?

In a durable database you control, keyed by tenant, user and thread, with clear ownership, retention and deletion rules. Not only in the browser or process memory, and not only in a provider's hosted thread feature unless you accept its retention and portability terms.

Language models are stateless: every request carries whatever context the model needs. So conversation state is your responsibility. It includes messages, tool calls and their results, intermediate agent steps, workflow checkpoints, attached files, summaries, and metadata such as the model and prompt version used for each turn.

Each storage option has a role. Browser storage is lost when the user changes device, and it can be edited, so a user could insert fake assistant turns or tool results; never treat client-supplied history as trusted. Server process memory disappears on restart and breaks when you run more than one instance. A cache such as Redis suits hot session data with a time-to-live, not the source of truth. Provider-hosted threads are convenient but tie history to one provider, follow its retention policy, and make fallback to another model harder. A relational database is the usual source of truth.

Keep messages append-only with a sequence number per thread. Store summaries and compressed context as separate derived records rather than overwriting the originals, so you can still replay and debug what the model actually saw.

Decide ownership and lifecycle before launch. Who can read a thread: the user, their organisation's admins, your support staff? How long is it kept? When a user deletes a conversation, deletion must reach every derived copy: summaries, embeddings, extracted memories, analytics exports and logs. Retrofitting that across five stores after launch is slow and error-prone.

Where should human approval go?

Immediately before actions that are consequential and hard to reverse, after deterministic validation, and wherever the measured error rate for a decision type exceeds the risk the business accepts. The human approves the exact action and parameters, not a plan.

Two properties decide whether an action needs approval: impact (money moved, data changed, messages sent to customers, legal or safety consequences) and reversibility (a draft can be discarded; a sent email or a paid invoice cannot). Add a third: the measured error rate of the system on that decision type from your evals and production review. A model's self-reported confidence is poorly calibrated and should not stand in for that measurement.

Position matters. The gate belongs after the model proposes an action and after code has validated it, and directly before execution. Validation catches what rules can catch: schema, amount limits, an approved supplier list, bank details unchanged since the last payment. Rule failures are rejected without bothering a person. Only actions that pass validation but sit above a risk threshold reach a reviewer, which keeps the queue small and meaningful.

Approval must bind to the exact payload. The reviewer sees the concrete action, the evidence behind it and what will change, and the approval record stores a hash of those parameters. If the agent alters the amount or recipient afterwards, the approval no longer matches and execution stops. Approvals should be recorded server-side with identity and time, never inferred from text in the conversation, because injected content can claim that a user already approved.

Design against approval fatigue. A reviewer who sees two hundred items a day starts approving without reading; an approval rate near 100% means either the gate is unnecessary or nobody is checking. Use tiers: automatic below a threshold, single approval in the middle, two-person approval at the top, and random sampling of automatic decisions for audit.

Approvals take minutes or days, so the run must checkpoint its state, wait without holding resources, resume on decision, and expire approvals that sit too long. Start with more gates than you think you need, measure how often reviewers change the outcome, and relax thresholds where the data says humans rarely intervene.

How should failed runs retry?

Retry only transient failures such as timeouts, rate limits and overloaded servers, with capped exponential backoff and jitter. Make every side-effecting step idempotent so a repeat cannot act twice, and send permanent errors straight to a fix or a human.

Start by classifying the error. Transient errors (network timeouts, HTTP 429 rate limits, 5xx server errors, provider overload) may succeed on a later attempt. Permanent errors (400 validation failures, 401 and 403 permission errors, a missing record, a business rule rejection such as an invalid account) will fail identically every time, so retrying only adds cost and delay. Model-quality failures, where output fails schema validation or a check, form a third category: retry once or twice with the validation error fed back, perhaps on a stronger model, then stop.

Back off exponentially with jitter (random variation) so thousands of clients do not retry in lockstep and flatten a recovering service. Honour Retry-After headers. Cap both the attempt count and the total time, and fit that budget inside the user's latency expectation for interactive paths.

A timeout does not mean the action failed. The payment may have gone through and only the response was lost. That is why every side-effecting step needs an idempotency key: a stable identifier for the logical action, such as run id plus step id, sent on every attempt so the downstream system performs it once. Where an API does not support keys, record the intent in your database before calling and check the target's state before retrying. Model calls have no side effects but cost money; tool calls are where duplicates hurt.

Retry the failed step, not the whole run. Checkpoint after each completed step so a crash resumes where it stopped. Rerunning an agent from the beginning repeats token spend and may take a different path, producing different side effects.

Watch for retry multiplication. An SDK that retries three times, inside your function that retries three times, inside a queue that redelivers three times, makes up to 27 attempts per message. Pick one layer to own retries. Add a system-wide retry budget and a circuit breaker so a degraded dependency fails fast instead of being hammered, and route exhausted runs to a dead-letter queue with full context for a person.

How should provider outages be handled?

Assume the provider will sometimes be slow or down. Set tight timeouts, detect degradation with a circuit breaker, and fail over to a pre-tested fallback that meets the same data policy and has passed your evals, down to a graceful reduced feature.

Outages rarely look like a clean failure. More often they are elevated error rates, latency spikes, rate limiting during regional peaks, one region degraded while others work, or a quiet behaviour change after a model update. Default client timeouts can be minutes long, which turns a slow provider into exhausted worker threads across your whole service. Set explicit timeouts per call: a connect timeout, a time-to-first-token timeout for streaming, and a total timeout sized to the expected output length.

Build a fallback ladder and know what each rung costs in quality. The usual order is the same model in another region or deployment (many providers and cloud platforms offer this), then an equivalent model from another provider, then a smaller model, then a reduced feature such as search results without a generated summary or a cached answer, then a clear message that the request is queued. Prompts often need adjustment for a different model, so each rung needs its own eval run, repeated on a schedule. A fallback that is never exercised quietly stops working.

A fallback must obey the same data policy as the primary: contracts and data processing terms, residency, retention and training restrictions. Sending regulated data to an unapproved provider during an outage is a second incident on top of the first. Tag requests with a data class and allow each rung only for the classes it is approved for.

Put the mechanics in one place, usually a gateway layer. It tracks health per provider and region, opens a circuit breaker when error rate or latency crosses a threshold, sends occasional probe requests to detect recovery, and limits retry traffic so recovery is not met with a retry storm. Latency-critical paths can use hedged requests, sending a second request if the first has not answered by the p95 time, at extra cost. Asynchronous work can simply queue and replay. Decide what happens to a stream that breaks mid-answer: restart on the fallback or show the partial text with a retry option.

Test it. Block the primary in staging and in scheduled drills, confirm traffic moves, measure quality on the fallback, and alert when any fallback has been active longer than expected.

How do I debug a bad answer?

Reproduce it from the trace, then walk the pipeline in order: input, retrieval, assembled prompt, model output, tool calls and results, post-processing. Find the first stage where something went wrong; most bad answers start before the model is called.

You cannot debug what you did not record. A useful trace holds the exact user input, any rewritten query, the retrieved candidates with document ids, versions and scores, the final assembled prompt including system text, history and tool definitions, the model id and parameters, the raw output, every tool call with arguments and results, and the post-processed answer the user saw. Log the prompt and configuration versions too, under the same access controls as the data itself.

Walk the stages in order and stop at the first one that is wrong. Was the request understood, or was it ambiguous, or did a query-rewriting step change its meaning? Was the right document among the retrieval candidates? If not, the problem is retrieval: chunking, the embedding model, a filter, or a stale index. Was it retrieved but dropped by reranking or by the context budget? Was it in the prompt but contradicted by another passage, such as last year's policy alongside this year's? Was it present and clear but ignored, which points at the prompt, the model or a very long context? Did a tool receive wrong arguments, return an error that was swallowed, or return data the model misread? Did a parser or the interface drop part of a correct answer?

Then reproduce. Replay the exact assembled prompt against the same model and settings several times. If it fails one time in five, the issue is variance and calls for a clearer instruction, examples, constrained output or a verification step. If it fails every time, change one variable at a time and rerun.

One complaint is an anecdote. Collect fifty failures, label each by the stage where it first went wrong, and fix the largest bucket. This error analysis often shows that retrieval or stale data causes more bad answers than the model does. Finally, add every fixed case to the regression eval set so the same failure cannot return unnoticed.

What must be checked before production?

Evidence, not a demo: task evals passing on representative cases, security boundaries holding against injection and cross-tenant access, cost and latency within budget at expected load, tested failure paths, and traces, dashboards and alerts in place. Then release gradually with rollback criteria.

Quality. An eval set drawn from real or realistic inputs, including edge cases and adversarial ones, with pass thresholds the product owner has agreed to. Scores reported per slice (language, customer tier, document type), not only overall. A regression suite that runs in CI on any prompt, model or retrieval change. A human read-through of a sample of outputs, because metrics miss things people notice in minutes.

Security and privacy. Prompt injection tests, including indirect injection through documents, web pages and tool results. Tool permissions enforced in code with least privilege, not by prompt instruction. Tenant isolation tests proving user A cannot retrieve user B's data. No secrets in prompts. PII handling, retention and the provider's data terms reviewed. Rendered output checked as an exfiltration channel: markdown images and links can leak data to an attacker's server.

Cost and performance. Cost per task measured at the median and the tail, because agent runs have long tails. Budgets, per-user rate limits and a spend alert. A load test at expected concurrency that includes the provider's own rate limits, not just your servers. Time to first token and end-to-end p95 latency within the target.

Failure handling. Timeouts on every external call, idempotent retries, a tested provider fallback, graceful degradation messages, step and spend limits for agents, approval gates on consequential actions, and a kill switch or feature flag that turns the feature off without a deploy.

Operations. Traces that record prompt and model versions, a dashboard for quality, latency, cost and errors, alerts with owners, a runbook, and a way for users to flag bad answers. Release in stages: shadow traffic, then a canary at a small percentage, with rollback criteria written down before launch so the decision is not argued during an incident.

How do you design from scratch?

Define the user outcome and how you will measure it, map the task as a workflow, make data, permissions and side effects explicit, build the simplest version with an eval set from day one, then add complexity only where evals and traces show a gap.

Start from the outcome, not the technology. "Resolve refund requests that meet policy without an agent touching them" is a design target; "build a support chatbot" is not. Define success as a number you can measure, such as resolution rate, handling time or error rate, along with what a wrong answer costs. That cost decides how much autonomy and review the design needs.

Map the task the way a skilled person does it today: inputs, decisions, lookups, actions and hand-offs. Mark each step as deterministic (code can do it), judgement (a model helps), or consequential (needs validation and possibly approval). Usually most steps are deterministic, a few need a model, and only one or two act on the world.

Make the hard-to-reverse decisions explicit early. What data does the system read, from where, and under whose permissions? What can it change, through which tools, with what idempotency and approval? Where does state live and how is it deleted? Which providers may see which data classes? These are expensive to change after launch; the model choice is not.

Build the simplest end-to-end version: one model, a workflow with model steps, whole documents or basic hybrid retrieval, Postgres for state. Collect 50 to 100 real cases before or during the build and run them on every change. Instrument traces from the first run.

Then iterate on evidence. Error analysis shows where the system fails; each fix targets the largest bucket. Add an agent step, a second model, a reranker or memory only when a failure category demands it, and confirm on the eval that it helped. Launch narrow, to one request type or one team, and widen as the numbers hold.

Should a request run synchronously or as a background job?

Keep it synchronous and streamed when the work finishes in seconds and the user waits for it. Make it a background job with durable state when it can run long, calls many tools, needs approval, or must survive a disconnect or deploy.

A synchronous request holds the connection open until the answer is ready, usually streaming tokens so the user sees progress. It is the simplest design and right for chat replies, autocomplete and short extractions. Its limits are physical: load balancers and gateways often cut idle or long connections after tens of seconds to a few minutes, mobile clients disconnect, deploys kill in-flight requests, and a failure halfway leaves nothing to resume.

A background job separates accepting the work from doing it. The API records a job, puts a message on a queue, and returns a job id at once. A worker processes it, checkpointing as it goes, and the client follows progress by polling, server-sent events or a notification. This brings retries, resumability, concurrency limits that protect provider rate limits, and natural pauses for human approval.

Agent runs, document batches and deep research belong in the background because their duration is unpredictable and they make many calls. A useful threshold is whether the p95 duration fits comfortably inside your shortest network timeout with margin to spare. Hybrids are common: stream a quick first answer synchronously while a job enriches it, or let a chat message start a long task and report back in the thread.

Background work costs more to build: a queue, workers, job state, progress UI and handling of redelivered messages. Some providers also offer batch interfaces that process large volumes of non-urgent requests over hours at reduced cost, worth checking for nightly workloads.

How do you keep the architecture adaptable as models change?

Put model calls behind one internal interface, keep model choice and prompts as versioned configuration, own your state and data, and keep an eval suite that can qualify a new model in hours. The eval suite is what makes switching safe.

Models change on a cycle of months: new versions arrive, old ones are deprecated with retirement dates, prices and rate limits shift, and behaviour changes in ways release notes do not fully describe. A system with a model name hard-coded in forty places, prompts tuned to one model's quirks and history stored only in a provider's thread feature is expensive to move. Adaptability is designed in, not bolted on.

Route every model call through one internal interface that covers what you use: messages, tool definitions, structured output, streaming and usage reporting. This can be a gateway or a thin adapter of your own. Avoid abstracting down to the lowest common denominator; expose provider-specific features such as prompt caching or reasoning effort as optional capabilities, so you can use them without coupling every caller to them.

Keep the model id, parameters and prompt version per task in configuration, and record them on every trace. Where providers offer pinned snapshots, pin them instead of using floating aliases, so behaviour does not change underneath you; upgrade deliberately after an eval run.

Own the data that is costly to rebuild. Conversation history and state belong in your database. Keep source text next to embeddings, because changing the embedding model means re-embedding the whole corpus, since vectors from different models are not comparable. Describe tools in a portable format; open protocols such as MCP help.

The decisive piece is the eval suite. When a new model appears or an old one is retired, run the suite per task slice, compare quality, cost and latency, and switch the tasks where the new model wins. With it, a migration takes days. Without it, every switch is a leap of faith. Also revisit scaffolding when models improve: retries, prompt chains and checks built to compensate for a weaker model may no longer earn their cost.