Understand where spend comes from and how to keep each useful outcome affordable.
Cost engineering is the practice of knowing what each useful outcome of an AI system costs, and keeping that number affordable without degrading quality. Unlike traditional software, where serving one more request is nearly free, every LLM request spends tokens, and the amount can vary by a factor of a hundred between a simple question and a long agent run. Spend that looks fine in a demo can sink a product's margin at scale, and a single stuck loop can produce a bill nobody expected.
The mental model has three layers. At the bottom is the token: input tokens you send, output tokens the model generates, with output priced higher because generation is sequential and memory-bound. Above that is the call and the run: context grows and is re-sent on every turn and agent step, so costs compound rather than add. At the top is the task and the customer: cost per successful task, cost per customer and gross margin are the numbers that connect engineering choices to business viability.
The questions follow those layers. The Basic questions explain how pricing works, what counts as input and output, and how long context, caching and agent loops change the bill. The Advanced questions cover the levers and controls: routing, budgets, batch inference, self-hosting economics, cost attribution, gross margin, surprise-bill prevention, and an ordered method for cutting cost without hurting quality.
Most providers bill per token, with separate rates for input tokens you send and output tokens the model generates, plus discounts for cached or batched input and extra charges for some tools, images, audio or very long contexts.
The core unit is the token, a chunk of text roughly three to four characters of English on average. A request's bill is usually input_tokens x input_rate + output_tokens x output_rate, with rates quoted per million tokens. Output is almost always priced higher than input, often several times higher, because generating tokens is the more expensive part of serving a model (see why outputs cost more than inputs).
On top of that base formula, most providers add modifiers. Cached input (a prompt prefix the provider has already processed recently) is billed at a steep discount, and some providers charge a small premium to write the cache. Batch or asynchronous endpoints, where you accept results within hours rather than seconds, are usually discounted. Some providers charge a higher tier once a single request crosses a long-context threshold. Reasoning tokens, the hidden thinking a reasoning model does before answering, are typically billed as output even when you never see them.
Non-text inputs are converted to tokens or billed per unit. Images are usually charged by resolution, so a full-size screenshot can cost as much as several pages of text. Audio may be billed per token or per minute. Built-in tools such as web search or code execution can carry a per-call or per-session fee on top of the tokens their results add to the context.
The practical consequence is that you cannot estimate cost from the prompt you typed. You need the token counts the API returns in its usage field for every call, broken into input, cached input, output and reasoning, and you need the provider's current rate card for each category. Tokenizers also differ between model families, so the same text can produce noticeably different token counts on two providers.
Input tokens are everything the model reads on a call: system prompt, tool definitions, conversation history, retrieved documents, tool results, images and the user's message. You pay for all of it on every call, not just the new text.
An input token is any token placed in the model's context window for a request. Users tend to think of input as the question they typed, but in a real application that is often the smallest part. A typical request also carries a system prompt with instructions and policy, JSON schemas for every tool the model may call, the conversation so far, passages retrieved from a knowledge base, and the raw results of earlier tool calls. Images, PDFs and audio are converted to tokens too.
The key fact for cost is that LLM APIs are stateless: the model does not remember the previous turn. To continue a conversation, your application resends the whole history each time. A ten-turn chat therefore pays for turn one ten times, turn two nine times, and so on, so total input grows roughly with the square of the number of turns. Agents are worse, because each tool result is appended and re-read on every later step.
Input is usually cheaper per token than output, but there is far more of it. In retrieval and agent workloads it is common for input to be 90% or more of all tokens and the majority of spend. That makes input the first place to look when a bill is high: oversized tool schemas, ten retrieved chunks when three would do, verbose tool results passed in unfiltered, and history that is never trimmed.
Two levers reduce input cost without changing what the model sees in substance. Prompt caching discounts a repeated prefix, so ordering stable content first matters. Context engineering reduces what you send: summarising old turns, filtering tool output to the fields the model needs, and retrieving fewer, better passages.
Output tokens are the tokens the model generates: the visible answer, any structured output or tool-call arguments, and on reasoning models the hidden thinking tokens. They are priced higher than input and also drive latency.
An output token is one the model produces. That includes the text the user sees, but also JSON for structured output, the arguments of every tool call an agent makes, and on reasoning models the internal reasoning tokens the model generates before its final answer. Most providers bill those reasoning tokens as output even if they are summarised or hidden, and the usage field reports them separately.
Output costs more per token than input, and it also costs time. The model generates tokens one at a time, so a 1,000-token answer takes roughly ten times as long to stream as a 100-token one. Reducing output therefore improves cost and latency together, which is rarely true of other optimisations.
Output length is partly under your control. Instructions such as "answer in at most three sentences" or a strict output schema shorten responses. The max_tokens parameter (named differently by some providers) sets a hard ceiling, but hitting it truncates mid-sentence or mid-JSON, so it is a safety cap, not a style control. Reasoning models usually expose an effort or thinking-budget setting, which is often the single largest output lever: high effort on a simple classification can generate thousands of reasoning tokens to produce a one-word label.
Watch for output that nobody reads. Models restate the question, add caveats, or produce a summary of what they just did. In agent pipelines, intermediate outputs consumed by code rather than people can be terse structured data. Each of those trims is small per call but compounds over millions of calls.
Cost scales with the tokens you send, so a long context raises the price of every call that carries it, and in multi-turn chats and agents that context is re-sent and re-billed on each step. Some providers also charge a higher rate above a size threshold.
Because input is billed per token, a 100,000-token context costs a hundred times more than a 1,000-token one on the same model. The model reads all of it whether or not it is relevant: there is no discount for passages it ignored. Some providers also apply a higher per-token rate once a request crosses a long-context threshold, so the curve can bend upward.
The bigger effect is repetition. A long document attached at the start of a chat is re-read on every turn, and an agent that pulls a large file into context pays for it on every later step. Ten steps over a 50,000-token context is 500,000 input tokens, even if each step only needed one paragraph. Prompt caching softens this when the long content sits in a stable prefix, but only within the cache lifetime and only for the exact repeated prefix.
Long context also costs in ways that do not show on the token line. Prefill time grows with length, so time to first token rises. On self-hosted models the KV cache (the stored attention keys and values for every token in context) consumes GPU memory per request, which reduces how many requests fit on a GPU at once and raises cost per request. And quality often drops as context grows, a pattern called context rot, so you may be paying more for a worse answer.
The alternatives are usually retrieval (send the relevant few thousand tokens instead of the whole corpus), summarisation of older material, and splitting work so each call sees only what it needs. Long context is the right choice when the task truly needs global view of a document, such as checking consistency across a contract, and the call is infrequent. It is the wrong choice as a default substitute for retrieval on a high-volume path.
Prompt caching lets a provider reuse the processed form of a prompt prefix it has seen recently, billing those tokens at a large discount and returning the first token faster. It only works when the start of the request is byte-for-byte identical.
When a model reads a prompt, it computes internal state (the KV cache) for every token. Prompt caching keeps that state for a recently seen prefix so a later request that starts with the same tokens can skip recomputing it. Providers pass the saving on: cached input tokens are billed at a fraction of the normal input rate, and time to first token drops because less prefill work is done.
The match is on the prefix, from the first token onward, and it must be exact. Change one character in the system prompt and everything after it is a cache miss. That is why the standard advice is to order context from most stable to least stable: system instructions and tool definitions first, then long reference documents, then conversation history, and the user's new message last. Putting a timestamp, request id or user name at the top of the system prompt defeats the cache for every call.
Details vary by provider, and you should check the current documentation. Some cache automatically when a prefix exceeds a minimum length (often around a thousand tokens); others require you to mark cache breakpoints explicitly. Some charge a premium to write an entry and a deep discount to read it. Entries expire after a short idle time, typically minutes, with longer lifetimes available on some platforms. The usage field reports cached tokens, which is how you confirm it is working.
Caching pays off when many calls share a long prefix within the cache lifetime: chat assistants with a large system prompt, agents that re-read the same history each step, and batch jobs that ask many questions about one document. It does little for short prompts or low traffic where entries expire before reuse. It is different from semantic caching, which stores whole responses and returns them for similar questions; prompt caching never skips the model call, only part of its input work.
Cost per task is the full cost of producing one finished, useful outcome: every model call, retry, embedding, retrieval, tool fee and slice of infrastructure it took, divided by tasks that actually succeeded. It is the unit that maps to business value.
Per-token and per-call costs are the wrong unit for decisions because users do not buy tokens. They buy a resolved ticket, a drafted contract, a merged pull request. Cost per task adds up everything spent to produce one such outcome: all LLM calls including routing, retries, judges and guardrail checks; embedding and vector search; paid tool and API calls; and an allocated share of infrastructure such as hosting, queues and observability.
The denominator matters as much as the numerator. If 100 attempts cost 50 units and 80 succeed, the cost per successful task is 0.625, not 0.5. Failed attempts, abandoned sessions and tasks a human had to redo are part of the cost of the ones that worked. Including human review time is often the biggest correction: an AI draft that costs 0.02 in tokens but takes an analyst ten minutes to fix is not cheap.
Cost per task enables comparisons that per-call metrics hide. A larger model may cost three times as much per call yet be cheaper per task if it succeeds first time where a small model needs retries and escalations. An agent with more steps may cost more per run but less per resolved case if its success rate is much higher. It also gives product and finance a number they can compare against what the task is worth or what it cost before.
Measuring it requires tagging every call, tool invocation and retry with a task id (usually the trace id) and recording the outcome. Report the distribution, not only the mean: agent tasks often have long tails where 5% of tasks consume 40% of spend.
Each loop iteration is another model call that re-reads the growing context, so agent cost grows faster than the number of steps. Retries, reflection passes and stuck loops multiply it further; a long run can cost hundreds of times a single answer.
An agent loop repeats: the model reads the context, chooses an action, a tool runs, and the result is appended. Every iteration is a full model call, and because the context grows with each tool result, every iteration is more expensive than the last. If step one reads 5,000 tokens and each step adds 2,000, step 20 reads 43,000 tokens and the run as a whole reads about 480,000. Cost grows roughly with the square of the step count, not linearly.
Several patterns inflate this. Retries after tool errors repeat full calls. Reflection or self-critique steps add calls that do not take actions. Planner and sub-agent designs spawn parallel model calls, each with its own context. And stuck loops, where an agent repeats the same failing search or edit, can burn budget for minutes with no progress. Large tool results, such as a whole file or a 500-row query result, make every later step heavier.
The useful levers are structural. Cap steps and tokens per run. Detect repetition (same tool with the same arguments twice) and stop or change strategy. Trim or summarise tool results before appending them, and compact history when it passes a threshold. Use caching so the stable prefix of each step is cheap. Route simple sub-steps to a smaller model. And ask whether the task needs an agent at all: a fixed workflow with two calls is often cheaper and more reliable than an open loop.
Agent cost is also highly variable. The median run may be cheap while a small fraction of runs consume most of the budget. Budget and monitor at the tail, not the average.
Routing sends each request to the cheapest model that can handle it well, using rules, a classifier or a cascade that escalates when a cheap model's answer fails a check. Savings come from the share of traffic that is easy, often the majority.
Price differences between model tiers within one provider are often an order of magnitude or more, and most production traffic is not hard. Greetings, simple lookups, formatting, extraction from clean text and short classifications rarely need the strongest model. Model routing exploits that: if 70% of requests can be served by a model costing a tenth as much, total spend falls by about 60% even though the hard 30% still go to the large model.
There are three common designs. Rule routing uses known signals: endpoint, task type, input length, customer tier. It is cheap, predictable and the right first step. Classifier routing runs a small model or a lightweight classifier to predict difficulty or intent, then picks a model. It adds a call, so the classifier must be much cheaper than what it saves. Cascades try the cheap model first, check the answer (schema validation, a confidence signal, a verifier or test), and escalate to the larger model only on failure. Cascades suit tasks with cheap, reliable checks, such as code that must compile or JSON that must validate.
The economics depend on three numbers: the share of traffic routed cheap, the quality loss on that share, and the overhead (classifier calls, or wasted cheap calls that escalate anyway). A cascade where 40% of requests escalate pays for the cheap attempt on those requests and adds latency. Routing also multiplies your evaluation burden: each route needs its own eval set, and a router misclassifying hard requests as easy produces quality failures that look random.
Routing should be evaluated like any model change. Build a labelled set covering easy and hard cases, measure each candidate model per segment, then pick thresholds that meet a quality floor at minimum cost. Monitor routed share and per-route quality in production, because traffic mix drifts and a router tuned last quarter can quietly send more traffic to the expensive path.
Enforce budgets in code at several levels: per call (max tokens), per run (steps, tokens, money), per user or tenant (quotas) and per organisation (spend caps and alerts). Check before each call, record after it, and stop or degrade gracefully when a limit is hit.
A budget that lives only on a dashboard is a report, not a control. Enforcement means the system refuses or changes behaviour when a limit is reached. It works best in layers, because each layer catches a different failure. Per call: max_tokens and a reasoning budget stop one runaway generation. Per run: limits on steps, tokens, wall time and money stop a stuck agent. Per user or tenant: daily or monthly quotas stop one heavy user or abusive account from consuming everyone's budget. Per organisation: provider-side spend limits and project keys cap the worst case if your own code fails.
The run-level budget is the one most teams lack. Implement it as a small object that travels with the task: before each model call, estimate the maximum cost of the call (input tokens plus max_tokens at output rates) and refuse if it would exceed what remains; after the call, record actual usage from the response. Pre-checking with the maximum, not the expected, cost prevents a single large call from overshooting the cap.
What happens at the limit is a product decision. Options include stopping with a partial result and a clear message, switching to a cheaper model, reducing reasoning effort, asking the user to confirm continuing, or queuing the task for later. Abruptly failing after spending most of the budget is the worst choice: you paid for work you then threw away. For agents, a useful pattern is a soft limit at 80% that tells the model to wrap up and summarise, and a hard limit at 100% enforced by code.
User and tenant quotas need shared state, usually a counter in a fast store keyed by user and period, updated atomically after each call. Pair them with rate limits on request count, because a quota measured in money only reacts after the spend. Finally, separate provider keys or projects per environment and per major feature, each with its own provider-side limit, so a bug in a batch job cannot exhaust the production budget.
Self-hosting is cheaper only when steady, high utilisation of the hardware brings cost per token below API prices after counting GPUs, idle time, engineers, monitoring and the quality gap of models you can run yourself. Spiky or modest traffic usually favours APIs.
An API charges per token, so you pay nothing when idle. A GPU charges per hour whether it serves one request or a thousand. The comparison therefore turns on utilisation: the fraction of the GPU's capacity you actually use. The cost per million tokens of a self-hosted model is roughly the hourly cost of the hardware divided by the tokens it produces per hour at your real load. Serving engines with continuous batching can process many requests at once, so throughput at high concurrency can be many times the throughput of a single stream, but only if traffic is there to fill the batch.
Hardware is the visible part. The full total cost of ownership also includes engineering time to deploy, tune, upgrade and patch the serving stack; on-call coverage; monitoring; capacity headroom for peaks (often 30 to 50% over average); redundancy across zones; and storage and networking. One or two engineers' time can exceed the GPU bill for a modest deployment.
Then there is quality. The model you can self-host may not match the frontier API models on your task, and a cheaper model that needs more retries, longer prompts or human correction can cost more per successful task. The comparison has to be on your eval set, at cost per successful task, not at cost per token.
Self-hosting tends to win in a few situations: high, steady volume (for example large nightly batch processing that keeps GPUs saturated); a small or fine-tuned model that does a narrow task well and fits on cheap hardware; strict data residency or air-gapped requirements where the API is not an option at any price; and latency-sensitive workloads that benefit from co-location. It tends to lose for spiky interactive traffic, early products whose volume is uncertain, and tasks that need the strongest available models. A hybrid is common: self-host a small model for a high-volume narrow task and use APIs for everything else.
Gross margin is revenue minus the direct cost of delivering the product, divided by revenue. For AI products that cost includes model inference, which scales with usage, so margin depends on how heavily each customer uses the product, not only how many customers you have.
Gross margin is (revenue - cost of goods sold) / revenue. Cost of goods sold (COGS) is the direct cost of delivering the service: model and API usage, embedding and vector storage, hosting, third-party data and tool fees, and often a share of support. A product earning 100 per month from a customer and spending 25 to serve them has a 75% gross margin. Traditional software usually runs at high margins because serving one more user costs almost nothing.
AI products differ because the marginal cost of usage is real and variable. Every request spends tokens, and heavy users spend far more than light ones. Under a flat subscription, revenue per user is fixed while cost per user follows usage, so the heaviest few percent of users can be served at a loss. Margin also moves with things outside your pricing: a feature that adds an agent loop, a prompt that grew by 3,000 tokens, or a switch to a reasoning model can drop margin by many points overnight.
Managing margin starts with measuring cost per customer and per feature, then comparing it to revenue per customer. The levers are on both sides. On cost: caching, routing, smaller models for easy work, shorter outputs, batch processing for non-urgent work. On pricing: usage tiers, credits, fair-use limits, charging for expensive features (long documents, deep research) separately, or outcome-based pricing where the price tracks the value delivered.
Do not over-index on today's number. Per-token prices for a given capability level have historically fallen, but teams usually spend those savings on more capable models, longer contexts and more agentic features, so cost per task does not fall automatically. Plan margin around the features you intend to ship, and model it at the usage of your heaviest customer segment, not the average.
Prevent surprise bills with hard limits at several layers, cost attribution on every call, near-real-time alerts on spend rate rather than monthly totals, and tests of worst-case paths such as retry storms, stuck agents and abusive users before they happen in production.
Surprise bills usually come from a short list of causes: an agent stuck in a loop, a retry policy that retries non-transient errors indefinitely, a batch job re-run by mistake over the full dataset, a prompt change that tripled context, a leaked API key, a single user or bot hammering an endpoint, and a model or effort setting changed in config. Each one is cheap to guard against and expensive to discover on the invoice.
Limits stop the damage. Per-call max_tokens, per-run step and money caps, per-user quotas and rate limits, retry policies with a maximum attempt count and backoff, and provider-side spending limits on each key or project. Use separate keys per environment and per major workload, so a test job cannot drain production budget and a leaked key has a small blast radius.
Visibility catches what limits miss. Compute cost per call from usage and attribute it to feature, model, user and environment. Alert on rate, not only on total: spend in the last hour compared with the same hour last week will page someone within an hour of a runaway loop, while a monthly budget alert fires weeks later. Add anomaly alerts on tokens per request and calls per task, which catch a prompt bloat or a loop even when traffic is normal.
Testing closes the gap. Before release, run a cost estimate over the eval set for any change to prompts, models, tools or retrieval, and compare to baseline. Deliberately trigger worst cases in staging: make a tool fail permanently and confirm the agent stops; send an enormous input and confirm it is rejected or truncated; simulate a thousand requests from one user and confirm the quota bites. Finally, have a kill switch per feature that ops can flip without a deploy.
Input tokens are processed in one parallel pass, while output tokens are generated one at a time, each needing a full pass through the model and holding GPU memory until the response ends. That sequential, memory-bound work is far more expensive per token, so providers price output higher.
Serving a request has two phases. In prefill, the model reads the whole prompt at once. All input tokens are processed in parallel as large matrix multiplications, which keep the GPU's compute units busy and produce many tokens of work per unit of time. In decode, the model generates the answer one token at a time. Each new token depends on the previous one, so each requires its own forward pass through every layer of the model.
Decode is limited by memory bandwidth, not compute. For every generated token the GPU must read the model's weights and the request's KV cache (the stored attention state for every earlier token) from memory, and do relatively little arithmetic with them. Batching many requests together amortises the weight reads, which is why servers batch decode steps across users, but the per-token cost of decode stays much higher than the per-token cost of prefill. A request also occupies KV cache memory for as long as it is generating, and that memory limits how many requests fit in a batch. Long outputs therefore tie up capacity that could serve other users.
Providers reflect this in pricing: output rates are commonly several times input rates, though the exact ratio varies by provider and model. Reasoning models amplify the effect, because their hidden thinking is also decode work billed as output. A response with 2,000 reasoning tokens and 200 visible tokens is priced on 2,200 output tokens.
The practical implications follow directly. In chat and generation tasks, output often dominates spend even though it is a small share of tokens. Cutting verbosity, using structured outputs, setting reasoning effort per task and avoiding unread explanations are among the highest-leverage savings, and they also reduce latency because decode time scales with output length. In retrieval and agent workloads input usually dominates instead, so measure the split before choosing what to optimise.
Use batch or asynchronous inference when results are not needed within seconds: evaluations, backfills, classification of a whole dataset, nightly reports. Providers commonly discount batch jobs substantially in exchange for completion within hours, and self-hosted servers run far more efficiently when batches are full.
Interactive serving has to keep capacity free for unpredictable bursts and return each answer quickly, which is expensive. Batch inference relaxes the latency requirement: you submit many requests together, usually as a file of independent requests, and collect results when the job finishes, often within a stated window such as 24 hours. The provider schedules the work into idle capacity, and many providers pass that efficiency on as a significant discount relative to their real-time rates. Check your provider's current terms, since the discount, completion window, size limits and supported features vary.
Good candidates share three properties: nobody is waiting on the answer, requests are independent, and the volume is large. Examples include running an eval suite on every release candidate, enriching or classifying a product catalogue, summarising yesterday's support tickets, generating synthetic training data, embedding a corpus for retrieval, and backfilling a new field across historical records. Batch discounts usually stack with prompt caching on some providers, which makes "many questions about the same document" jobs especially cheap.
Batch is a poor fit for anything user-facing, for multi-step agents where each step depends on the previous result (each step would wait for a batch cycle), and for jobs where a late result is worthless. A middle ground is an internal queue: interactive features enqueue non-urgent work, such as generating a weekly digest, and a worker drains the queue through the batch endpoint.
Engineering a batch pipeline well means handling partial failure. Jobs can complete with some requests errored or expired; you need per-request ids, a results reconciler that retries only failures, and idempotent writes so a re-run does not duplicate records. Validate a small sample synchronously before submitting a million requests, since a prompt bug in a batch job wastes the whole job.
Tag every model call with feature, user, tenant, model and trace id, compute its cost from the returned usage, and store it with the trace. Then roll up by any dimension. Without per-call attribution, a monthly invoice cannot tell you what to fix or whom to charge.
A provider invoice shows spend per model and maybe per API key. It cannot tell you that the document-summary feature costs four times the chat feature, or that one enterprise tenant accounts for a third of spend. Cost attribution answers those questions by recording, for every model call, who and what it was for, and what it cost.
The mechanism is simple and best put in one place: the wrapper or gateway every model call passes through. Each call carries metadata such as feature, tenant_id, user_id, environment, prompt_version and the trace_id of the task it belongs to. When the response returns, the wrapper reads the usage (input, cached input, output, reasoning tokens), multiplies by the rate table for that model, and emits a record. If you already use tracing, add the cost as an attribute on the LLM span so the same data serves debugging and finance. Non-LLM costs, such as embeddings, vector queries and paid tools, should be tagged the same way.
Once records exist, the questions become queries. Cost per feature shows where engineering effort pays off. Cost per tenant compared with revenue per tenant shows margin by customer and supports usage-based pricing. Cost per prompt version catches a regression the day it ships. Cost per trace, joined to outcome, gives cost per successful task. Shared costs that cannot be tagged per call, such as fixed hosting, can be allocated by traffic share, but keep them visibly separate from directly measured costs.
Two cautions. Recompute historical cost when rates change only if you keep the rate table versioned; otherwise store the computed cost at call time. And treat user ids in cost records as personal data under your retention and privacy rules, the same as any other telemetry.
Measure first, then apply levers in order of safety: remove waste (unused context, verbose output, broken caching), use discounts (caching, batch), then substitute (smaller models, routing, lower reasoning effort), checking each change against an eval set and cost per successful task.
Cost cutting goes wrong when it starts with the most visible lever, usually swapping to a cheaper model, and finds out about the quality loss from users. A safer approach orders changes by how much they risk quality, and gates each one on the same eval set used for any other change. The metric to optimise is cost per successful task, because a cheaper call that fails more often can raise total cost.
The first tier is waste removal, which rarely affects quality and often improves it. Trim tool schemas and retrieved passages to what the task needs; filter tool results; summarise old history; remove instructions the model ignores; shorten verbose outputs and drop unread explanations. Fix prompt layout so the stable prefix is cacheable. Stop retrying errors that will never succeed. Teams commonly find that a large share of tokens fall in this category.
The second tier is discounts that leave behaviour unchanged: prompt caching for repeated prefixes and batch endpoints for non-urgent work. These change the price of the same tokens, so quality risk is near zero, though batch changes latency.
The third tier is substitution, which does carry quality risk and must be evaluated per segment: lowering reasoning effort, routing easy traffic to smaller models, distilling or fine-tuning a small model for a narrow high-volume task, semantic caching of whole answers for repeated questions, and replacing an open agent loop with a fixed workflow. Each of these can save a lot, and each can fail on a slice of traffic that an aggregate score hides. Roll them out behind a flag, compare cost and quality on live traffic, and keep the ability to revert.