The GenAI Field Guide

Production architecture

Put a dependable application boundary around model calls and data.

A model call is the easy part of a GenAI product. The hard part is everything around it: who is allowed to make the call, how much it may cost, what happens when the provider is slow or down, where the conversation and the documents live, how a forty-minute agent run survives a deploy, and how ten thousand users share a rate limit that the provider sets per account. Production architecture is the set of components and boundaries that make model calls dependable, affordable and safe at the scale of real traffic.

The mental model is a conventional web backend with one unusual dependency. The model behaves like a slow, expensive, rate-limited, occasionally unavailable third-party API whose output is non-deterministic and whose cost depends on the size of the request. Almost every pattern in this chapter is a classic distributed-systems pattern (a backend boundary, a gateway, rate limits, queues, durable state, caching, idempotency, horizontal scaling) adjusted for that dependency: limits counted in tokens instead of requests, timeouts measured in tens of seconds, responses that stream, and caches that must respect meaning and permissions.

The Basic questions walk through the components in the order a request meets them: the backend boundary, the gateway, streaming back to the user, rate limits, queues for background work, and the two stores almost every system needs, PostgreSQL and Redis. The Advanced questions cover the decisions that separate a demo from a service: when a dedicated vector database earns its place, how provider fallback and semantic caching go wrong, how to build agents that run for minutes or hours, how scaling and concurrency control actually work when the bottleneck is someone else's GPU, and how to set timeouts, retries and versioned configuration so that changes and failures are both survivable.

What does a production AI architecture include?

A production AI architecture wraps model calls in a normal backend: a client, an API layer with auth and quotas, a model gateway, data stores, retrieval, tools, background workers, observability, evaluation and security controls, each with a clear owner.

A prototype is usually a script that sends a prompt to a model and prints the answer. A production system has to answer harder questions: who is this user and what may they see, how much can they spend, what happens when the provider returns errors for ten minutes, how do we know answers got worse after Tuesday's prompt change, and how do we delete a customer's data on request. Each of those questions maps to a component.

The request path is the client (web, mobile, chat integration), an API backend that authenticates the user and enforces quotas, an orchestration layer that builds the context (retrieval, conversation history, tool definitions) and calls the model through a gateway, and a response path that streams tokens back. The data layer holds durable records (users, conversations, documents, jobs, audit logs), a search index for retrieval, and a cache. The async layer is a queue and workers for ingestion, long agent runs and batch jobs.

Around the path sit the cross-cutting concerns. Observability records a trace per request with prompt version, model, token counts, latency and cost. Evaluation runs offline test sets before a change ships and samples live traffic afterwards. Security covers secret management, least-privilege tool credentials, input and output guardrails, and tenant isolation. Configuration versions prompts, model choices and feature flags so changes can be rolled out and rolled back.

Not every product needs every box on day one. A small internal assistant can run as one backend service, one PostgreSQL database and a hosted model. What should exist from the start are the boundaries: model calls go through one module you own, every call is logged with its cost, and state lives in a durable store. Those three make it cheap to add the rest later.

Should the frontend call a model directly?

Almost never. The browser or mobile app should call your backend, which holds the provider key, authenticates the user, enforces quotas and policy, builds the prompt and logs the call; anything shipped to a client can be extracted and abused.

Any secret in a frontend bundle or mobile binary is public. Someone opens developer tools, copies the API key and spends your budget on their own workload. Even without a leak, a client that talks straight to the model can send any prompt it likes, so your system prompt, tool definitions and usage limits are merely suggestions. The backend is where you can enforce rules, because the user cannot change code that runs on your servers.

The backend also owns the parts of the request that should not be visible or editable: the system prompt, retrieved documents filtered by the user's permissions, tool credentials, and the choice of model. It can check a per-user or per-tenant quota before spending money, redact personal data, attach a trace id, record token usage for billing, and apply output checks before anything is shown. Moving prompt construction server-side also lets you change prompts and models without shipping a new app version.

There are narrow exceptions. Some providers issue ephemeral, scoped tokens for real-time use cases such as voice sessions, where the client connects directly for latency reasons; the backend still mints the token, sets its lifetime and limits, and decides what the session may do. On-device models are another case, since nothing leaves the device. In both, the principle holds: the backend decides who may call what, with which limits.

The cost of the backend hop is small. It adds a few milliseconds compared with model latency measured in hundreds of milliseconds to many seconds, and streaming through the backend keeps time to first token nearly unchanged.

What is an AI gateway?

An AI gateway is a shared service, or library, that every model call passes through. It gives one place for provider credentials, routing, retries, fallback, rate limits, budgets, logging, caching and policy checks across teams and providers.

Once more than one team or feature calls models, the same concerns appear everywhere: holding keys, handling a 429 from the provider, tracking spend per team, switching a model, redacting personal data, logging prompts for debugging. An AI gateway centralises them. Applications call the gateway with a logical request (a model alias, messages, tools, a tenant id) and the gateway handles the provider-specific details.

Typical gateway features are a unified API over several providers, model aliases so applications ask for support-default rather than a vendor model id, retries and fallback with circuit breakers, rate limiting and budgets per team, tenant or key, usage accounting in tokens and estimated cost, logging and tracing with consistent metadata, optional caching, and policy hooks such as PII redaction, blocked topics or allowed-model lists per data classification. Security-focused gateways add prompt-injection screening and output filtering.

A gateway can be a separate network service (open-source proxies, managed products from cloud providers, or your own) or an in-process library shared by services. A service gives central control and works across languages, at the cost of an extra hop and a component that must be highly available, because if the gateway is down every AI feature is down. A library avoids the hop but needs every service to upgrade to get a policy change.

The main trade-off is abstraction versus capability. Providers differ in tool-calling formats, structured output support, caching controls and multimodal inputs. A gateway that normalises everything to a lowest common denominator hides features you may want. Good gateways normalise the common path and let callers pass provider-specific options through.

What is streaming?

Streaming sends the model's output to the user in small pieces as tokens are generated, instead of waiting for the full response. It cuts perceived latency from the full generation time to the time to first token.

A model generates text one token at a time. A 600-token answer at a decode speed of 50 tokens per second takes about 12 seconds to finish, but the first token may be ready in under a second. Without streaming the user stares at a spinner for the whole 12 seconds. With streaming they start reading almost immediately, and because people read slower than most models generate, the answer often feels instant. The total time does not change; the perceived latency does.

Providers stream over HTTP, most commonly with server-sent events (SSE), a simple format where the server keeps the response open and writes data: lines as chunks arrive. Your backend reads the provider stream, optionally transforms or checks it, and forwards it to the browser over SSE, a WebSocket, or a chunked fetch response. Tool calls, usage counts and finish reasons also arrive as stream events, often at the end.

Streaming complicates several things. Output guardrails that need the whole answer cannot run before the user sees the first words, so teams either check chunks incrementally, buffer a sentence at a time, or accept a post-hoc retraction. Structured output such as JSON is incomplete until the end, so parse it with a tolerant parser or only render it at completion. Infrastructure must not buffer: reverse proxies and some serverless platforms buffer responses by default, which silently turns a stream back into one late blob. And client disconnects should cancel the upstream request, otherwise you keep paying for tokens nobody will read.

Stream anything a person waits for. Do not bother for background jobs, batch processing or calls whose output is consumed by code that needs the complete result.

What is rate limiting?

Rate limiting caps how much work a caller may request in a time window. For AI systems it protects shared provider quotas and budgets, and it should count tokens and concurrent requests as well as requests per minute.

Rate limits exist at two levels. Providers limit your account, usually in requests per minute and tokens per minute (input, output, or both), and return HTTP 429 with a retry hint when you exceed them. Your application should limit its own users and tenants, so that one customer running a script cannot use the whole provider quota and starve everyone else, and so that a bug or abuse cannot produce a five-figure bill overnight.

Requests per minute alone is a weak limit for LLM workloads. One request with a 100,000-token document costs as much as hundreds of short chat turns. Good designs combine three limits: a request rate to stop floods, a token budget per minute or per day to bound spend and provider usage, and a concurrency limit on in-flight requests, because a long generation holds capacity for many seconds. Token counts are not known exactly until the response ends, so reserve an estimate up front (input tokens plus the max_tokens cap) and reconcile with the actual usage afterwards.

The usual algorithms are the token bucket, which refills at a steady rate and allows short bursts up to the bucket size, and the sliding window, which counts usage in the last N seconds. Both need a shared, atomic counter when you run more than one server, which is why Redis is the common choice. When a limit is hit, return 429 with a Retry-After header and a clear message in the UI, rather than queueing silently until timeouts fire.

Set your internal limits below the provider's, leaving headroom for retries and background jobs, and alert when aggregate usage approaches the provider ceiling.

What is a queue?

A queue is a durable buffer between the code that asks for work and the workers that do it. It lets slow or bursty AI tasks such as ingestion, long agent runs and batch generation happen in the background, at a controlled pace, with retries.

A web request should finish in seconds. Parsing a 300-page PDF, embedding it, and summarising each section can take minutes, and running it inside the request means timeouts, lost work on deploys and a user staring at a spinner. With a queue, the API writes a job record, enqueues a message, and returns a job id immediately. Workers pull jobs, process them, and update the status, which the client polls or receives through a push channel.

Queues also smooth bursts. If 2,000 documents arrive at once, the queue holds them and a fixed pool of workers processes them at a rate the provider's rate limit can sustain. Without the queue, the burst becomes 2,000 simultaneous model calls and a wall of 429 errors. And queues give retries: if a worker crashes or a call fails, the message becomes visible again and another worker picks it up, with a dead-letter queue collecting jobs that fail repeatedly so someone can look at them.

Most queues deliver at least once, so a job can run twice (for example when a worker finishes but crashes before acknowledging). Handlers must be idempotent: writing results with an upsert keyed on the job id, and guarding side effects such as sending an email with an idempotency key. Long jobs need a visibility timeout or lease longer than the expected run time, with heartbeats to extend it, otherwise the queue hands the same job to a second worker while the first is still working.

You do not need a dedicated broker on day one. A PostgreSQL table with SELECT ... FOR UPDATE SKIP LOCKED is a reliable queue for thousands of jobs per minute and keeps job state transactional with your data. Move to a managed queue or broker when throughput, fan-out to many consumers, or cross-service events demand it.

What belongs in PostgreSQL?

PostgreSQL holds the durable truth: users, tenants, permissions, conversations and messages, documents and their chunks, jobs and their status, audit logs, usage records, and at moderate scale the embeddings themselves through pgvector.

The rule of thumb is simple: anything you would be upset to lose, anything that needs transactions, and anything you will query by relationships belongs in a relational database. PostgreSQL is the common default because it is mature, widely hosted, and its extensions cover AI needs: pgvector for embeddings and approximate nearest-neighbour search, full-text search for keyword retrieval, and jsonb for semi-structured payloads such as tool results.

A typical AI product stores identity and access (users, tenants, roles, document permissions), conversation state (conversations, messages with role, content, model, prompt version and token counts), knowledge (documents, chunks, embeddings, source metadata and ingestion version), work (jobs, agent runs, checkpoints, tool calls with arguments and results), and accountability (audit logs of actions, usage and cost per tenant, user feedback). Keeping these together means a single transaction can save an agent's checkpoint and mark a job step done, and a single query can retrieve chunks filtered by the user's permissions.

Keeping embeddings in the same database as permissions is a real advantage. A retrieval query can join chunk vectors to an access-control table, or rely on row-level security, so a user never retrieves text they are not allowed to see. With a separate vector store you must copy permission metadata into it and keep it in sync, which is a common source of leaks.

PostgreSQL is not the right home for everything. Rate-limit counters that change on every request, short-lived caches and pub/sub signals fit Redis better. Large binaries such as PDFs and audio belong in object storage with a reference in the database. And very large or very high-query-rate vector workloads may outgrow pgvector, which is a measured decision, not a default.

What belongs in Redis?

Redis holds fast, short-lived, rebuildable state: rate-limit and quota counters, response and embedding caches, session data, distributed locks and semaphores, and lightweight queue or pub/sub signalling. If losing it would lose real data, it belongs in PostgreSQL instead.

Redis is an in-memory data store with atomic operations and expiry on every key. That combination fits the hot, ephemeral state an AI backend touches on every request. Counters for rate limits and token budgets need atomic increments shared across many servers. Caches for embeddings of frequent queries, retrieval results or full responses need sub-millisecond reads and automatic expiry. Coordination such as a distributed lock to stop two workers re-indexing the same document, or a semaphore capping a tenant's concurrent model calls, needs atomic check-and-set.

Redis is also commonly used for streaming fan-out: a worker generating a long answer publishes chunks to a channel or stream, and whichever web server holds the user's connection relays them. Redis Streams and list-based queue libraries are reasonable for simple background jobs, though jobs that represent real work should still have a durable record in the database.

The defining property is that Redis data should be disposable. Even with persistence enabled, Redis is usually configured with an eviction policy that drops keys under memory pressure, and failover can lose the most recent writes. So design every use so that a missing key causes a cache miss, a recomputed count or a retried lock, never a lost conversation or a forgotten payment. A useful test is to imagine flushing Redis during peak traffic: the system should get slower, not wrong.

Two cautions specific to AI workloads. Cached responses can contain personal or tenant data, so keys must include tenant and permission scope and values need a time to live. And large cached values such as long documents or big embedding batches can fill memory quickly; set size limits and monitor memory and eviction rates.

When do you need a vector database?

You need a dedicated vector database when measured needs exceed what your existing database with a vector extension can serve: very large collections, high query rates at strict latency, heavy filtered search, or features your database lacks. Most products should start without one.

A vector database stores embeddings and answers approximate nearest-neighbour queries quickly, usually with graph-based indexes such as HNSW or partition-based indexes such as IVF. The question is not whether you need vector search (most RAG systems do) but whether it needs its own system. PostgreSQL with pgvector, and the vector features of several search engines and document databases, provide the same core capability inside a store you already run.

Staying in your main database has concrete advantages. Vectors sit next to the documents, permissions and tenant ids, so a filtered search is one SQL query and access control is enforced by the same mechanism as the rest of the app. Writes are transactional, so a document and its chunks appear and disappear together. There is one backup, one security review and one on-call rotation. For collections in the low millions of vectors and moderate query rates, a well-tuned HNSW index in PostgreSQL usually returns results in tens of milliseconds, which is small next to model latency.

Dedicated systems earn their place in specific situations. Scale: hundreds of millions or billions of vectors, where index build time, memory for HNSW graphs, and sharding become the main problems. Throughput: thousands of queries per second with tight latency targets, where a purpose-built engine and horizontal sharding help. Filtered search at scale: selective metadata filters can wreck recall or latency in a naive approximate index, and some dedicated engines handle filtering inside the index traversal. Features: built-in hybrid search, multi-vector or sparse-vector support, tiered storage, or per-tenant namespaces. The thresholds vary a lot with vector dimension, hardware and filter patterns, so treat any number as a rough guide.

Decide with a benchmark on your own data: load a realistic corpus, run your real queries with your real filters, and measure recall at K against exact search together with p95 latency. Many teams discover the bottleneck is chunking or reranking quality rather than the index.

What is provider fallback?

Provider fallback routes a request to a different model or provider when the primary fails, times out or is rate limited. It helps only if the fallback is tested on your prompts, approved for your data, and triggered by a circuit breaker rather than every error.

Model providers have outages, partial degradations, regional incidents and capacity limits. A fallback chain lists alternatives in order: the same model in another region or deployment, a different model from the same provider, then a model from another provider, and finally a non-model degraded mode such as a canned message or search results without a generated answer. The gateway tries the next entry when the current one returns a retryable error or exceeds its deadline.

The hard part is that a fallback model is a different product. Prompts tuned for one model can behave differently on another: tool calls may be formatted or chosen differently, structured output may be less reliable, refusals and tone differ, and context windows and token counting vary. A fallback that has never been run against your eval set may turn an outage into a quality incident that is harder to notice. Run the eval suite on every model in the chain, keep per-model prompt variants where needed, and decide per feature whether a fallback is acceptable at all.

Fallback also crosses data boundaries. If the primary is approved under a data processing agreement with specific residency and retention terms, the fallback provider must be approved for the same data classification. Many teams allow cross-provider fallback only for low-risk features.

Trigger fallback with a circuit breaker: after a threshold of failures or timeouts within a window, open the circuit and send traffic straight to the fallback for a cooldown period, then probe the primary with a small share of requests before closing the circuit again. This avoids paying the primary's timeout on every request during an outage. Distinguish error types: a 400 for an invalid request should never trigger fallback, because the alternative will reject it too.

What is semantic caching?

Semantic caching returns a stored response when a new request is close enough in meaning to an earlier one, measured by embedding similarity. It can cut cost and latency for repetitive, non-personal questions, but a wrong hit serves a confidently wrong or leaked answer.

An exact cache only hits when the request text matches byte for byte, which rarely happens with natural language. A semantic cache embeds the incoming query, searches a store of previous queries, and if the nearest one has similarity above a threshold, returns its cached answer without calling the model. 'How do I reset my password?' and 'I forgot my password, how do I change it?' can share one answer.

The danger is that similarity is not equivalence. 'How do I cancel my order?' and 'How do I cancel my subscription?' may embed very close together and need different answers. Small words carry meaning that embeddings can underweight: negations, numbers, dates, product names, 'not'. A threshold loose enough to give a useful hit rate will produce some wrong hits, and the user cannot tell, because the cached answer looks fluent. Tune the threshold on labelled pairs of should-match and should-not-match queries from your own traffic, and measure the false-hit rate, not just the hit rate.

Scope is the second danger. A cached answer depends on everything that went into it: the system prompt version, the model, the retrieved documents, the user's permissions and the tenant. The cache key must include those, or at least tenant, permission scope and prompt version, otherwise one customer can receive an answer built from another customer's documents. Answers that depend on personal or live data such as an account balance, an order status or anything with a date should not be semantically cached at all.

Do not confuse this with prompt caching, which providers offer to reuse the computation of a repeated prompt prefix. Prompt caching still generates a fresh answer and is safe by construction; semantic caching skips generation. Many teams get most of the savings from prompt caching, exact caching of normalised queries, and caching retrieval results, and reserve semantic caching for public FAQ-style traffic with a high repeat rate.

How should long-running agents be built?

Run them as durable background jobs, not inside a web request: persist state after every step, make each tool call idempotent, enforce step, time and cost budgets, report progress, and support pausing for human approval and resuming after crashes or deploys.

An agent that researches for 40 minutes, makes 120 model calls and touches external systems will not survive in an HTTP request. Load balancers time out, deploys restart processes, workers crash, and providers return errors midway. The design goal is that any interruption costs at most one step of work, never the whole run, and never a duplicated side effect.

The core pattern is a loop over persisted state. The run lives in the database: goal, plan, completed steps with their outputs, pending step, budget consumed, status. A worker claims the run with a lease, loads state, executes one step (a model call or a tool call), writes the result and the next state in one transaction, and repeats. If the worker dies, the lease expires and another worker resumes from the last checkpoint. Durable execution engines formalise this by recording each step's result and replaying the workflow code deterministically; a hand-built version with a run table and step records works well for simpler agents.

Side effects need idempotency keys. If the agent sends an email or creates a ticket, derive a key from the run id and step number and pass it to the downstream API, or record the intent before the call and check for it before retrying. Otherwise a crash between 'action done' and 'checkpoint saved' repeats the action on resume.

Long runs also need limits and visibility. Enforce a maximum number of steps, wall-clock time and token or cost budget, and fail the run cleanly with a partial result when one is reached. Publish progress events (step started, tool called, waiting for approval) so the UI can show status and the user can cancel. Model human approval as a state: the run writes 'awaiting approval', releases its worker, and a later approval event re-enqueues it, so nothing holds a process open for hours.

Finally, plan for change. A run started on Monday with prompt v7 may resume on Tuesday after v8 ships. Store the configuration version on the run and either finish it on the old version or define how state migrates.

How does horizontal scaling work?

Horizontal scaling adds more identical, stateless instances behind a load balancer or queue, with all state kept in shared durable stores. For AI systems, the ceiling is usually provider rate limits, database connections and per-tenant fairness rather than your own CPU.

To scale horizontally, any instance must be able to serve any request. That means no conversation history, job progress or rate-limit counters held in process memory. State goes to PostgreSQL, Redis or object storage, and instances become interchangeable. Then you add instances behind a load balancer for API traffic, or add workers consuming a queue for background work, and capacity grows roughly linearly until a shared dependency saturates.

AI workloads have an unusual profile. A web server waiting on a model call uses almost no CPU but holds a connection and some memory for 5 to 60 seconds. Scaling on CPU therefore gives the wrong signal; scale on in-flight requests, queue depth or queue age instead. Use asynchronous I/O so one instance can hold hundreds of open streams. Long-lived streaming connections also make deploys and scale-in harder: drain connections gracefully, and design clients to reconnect and resume from a known message.

The real ceilings are elsewhere. Provider quotas: if your account allows a fixed number of tokens per minute, doubling workers doubles 429 errors, not throughput. Concurrency must be capped globally, and raising the ceiling means a higher quota, more deployments or regions, or batch APIs for non-urgent work. Database connections: fifty instances with a pool of twenty each is a thousand connections, more than many PostgreSQL servers handle well, so add a connection pooler and size pools deliberately. Vector search and embedding throughput may need their own replicas. If you self-host models, GPU capacity is the expensive, slow-to-scale tier, and batching efficiency matters more than instance count.

Finally, more capacity does not guarantee fairness. Without per-tenant limits, one large customer's batch import fills every worker and the rest wait. Separate queues or weighted scheduling by priority and tenant keep interactive traffic responsive while bulk work proceeds.

What is concurrency control?

Concurrency control limits how many operations run at the same time and prevents conflicting or duplicate work. In AI systems it means capping in-flight model calls globally and per tenant, and using locks, leases or optimistic versioning so two workers never process or edit the same thing at once.

Rate limits cap work per unit of time; concurrency limits cap work in flight at any moment. The difference matters for LLM calls, because a single generation can hold capacity for 30 seconds or more. If a tenant may make 60 requests per minute but each takes 40 seconds, they can have 40 calls running simultaneously, which may exceed the provider's concurrent request limit or your worker pool. A semaphore per tenant, per feature and globally keeps the number of in-flight calls bounded. Across many instances it must be distributed, typically a Redis sorted set of leases with expiry, so a crashed holder does not leak a permit forever.

The second job is preventing duplicate and conflicting work. Users double-click send; queues redeliver; two webhook events arrive for the same document. Without control, the system generates two answers, re-indexes a document twice concurrently, or lets two agent steps edit the same record. The tools are classic: a unique constraint or idempotency key so the second insert fails cleanly; a lock or lease (a Redis key with expiry, or a PostgreSQL advisory lock) so only one worker processes a given document or run; and optimistic concurrency with a version column, where an update succeeds only if the version is unchanged, which suits agents editing shared state.

Conversation turns need ordering too. If a user sends two messages quickly, should both be answered in parallel against the same history, or should the second wait? Most chat products serialise turns per conversation, either with a per-conversation lock or by rejecting a new turn while one is generating.

Concurrency limits also enable load shedding. When all permits are taken, decide explicitly: queue with a bounded wait, return 429 with a retry hint, or degrade to a cheaper path. An unbounded wait just converts overload into timeouts.

How should timeouts and retries be set for model calls?

Give every model call a deadline that fits the user's wait, retry only transient errors with exponential backoff and jitter, cap retries with a budget, honour retry hints, and use idle timeouts for streams. Never blindly retry calls that trigger side effects.

Model calls fail in ways ordinary APIs rarely do. A request can be rate limited (429), hit an overloaded or failing backend (5xx), or simply take far longer than usual because the output is long or the provider is degraded. The default timeouts in HTTP clients are either too short for a long generation or so long that a stuck request holds a connection for minutes. Without a deliberate policy you get either premature failures or a pile-up of hung requests during an incident.

Start from the deadline: the total time the caller can wait. An interactive chat answer might allow 60 seconds end to end; a background summary might allow 5 minutes. Every attempt and every backoff must fit inside it, so pass the remaining time down rather than giving each layer its own fixed timeout. For streams, use two timeouts: one for the first token and an idle timeout between chunks, so a stream that stalls midway is detected in seconds rather than at the overall deadline. Size max_tokens to the task so that a runaway generation cannot exceed the deadline.

Retry only transient errors: timeouts, connection resets, 429 and most 5xx responses. Do not retry 400-class validation, authentication or context-length errors, because they will fail again. Use exponential backoff with jitter (random spread) so many clients do not retry in lockstep, and prefer the provider's Retry-After value when it is present. Cap attempts, and add a retry budget that limits retries to a fraction of total traffic, so a provider outage does not triple your load on it.

Retries interact with everything else. A retried call that already started streaming to the user should not restart silently; show the failure or resume visibly. A retried call that triggered a tool action needs an idempotency key. And retries in several layers multiply: three retries in the SDK, three in the gateway and three in the job runner can make 27 attempts. Decide which layer owns retries and disable them elsewhere.

How do you version and roll out prompt and model changes?

Treat prompts, model choices and generation settings as versioned configuration: store them in source control or a registry, pin model versions, record the version on every trace, gate changes on evals, and roll out behind flags with instant rollback.

In an AI system, behaviour changes when the code changes, but also when someone edits a prompt, swaps a model, changes the temperature, updates a tool description, or the provider updates the model behind an alias. If those changes are not versioned, a quality drop is almost impossible to diagnose: nobody can say what was different on Tuesday. The fix is to give the full generation configuration an identity.

A configuration version bundles everything that shapes output: the prompt templates, the model identifier, sampling parameters, max_tokens, tool definitions, retrieval settings such as top K and reranker, and guardrail settings. Store it in source control or a prompt registry, review changes like code, and give each version an id. Pin model versions where the provider offers dated or snapshot identifiers, because a floating alias can change underneath you; when you do move, do it as a deliberate configuration change with an eval run. Record the configuration id on every trace and stored message, so any response can be traced to the exact setup that produced it.

Rollout follows the same path as a code release. Run the offline regression eval for the new version and compare it with the current one. Ship it behind a feature flag or traffic split, start with a small share of users or an internal cohort, watch quality, cost and latency signals, then widen. Keep the previous version deployable, so rollback is a flag flip rather than a code revert and redeploy. Prompts loaded at runtime from a registry make this fast, but they also make it easy to bypass review, so the registry needs the same approvals and audit trail as code.

Separate what can change without a deploy from what cannot. Many teams let product or domain experts edit prompts through a registry with eval gates, while model swaps, tool schemas and guardrail changes go through engineering review because they can break parsing or security assumptions.