Make responses feel fast without losing quality where it matters.
A GenAI feature that gives the right answer in 14 seconds often loses to one that gives a slightly worse answer in 3. Users judge speed by when something useful appears and when the job is done, and LLM systems are slow in ways that ordinary web services are not: generation is sequential, one token at a time, and a single request can fan out into retrieval, reranking, several model calls and a chain of tool calls. Performance engineering is the discipline of finding where that time goes and removing it without quietly removing quality.
The mental model is a latency budget split into two very different kinds of work. Prefill reads the whole input in parallel and decides how long the user waits for the first token. Decode produces output one token at a time and decides how long the stream lasts. Around those sit the non-model costs: network, queueing, retrieval, tool calls and orchestration. Each lever affects a different part: shorter prompts and prompt caching cut prefill, shorter outputs and faster models cut decode, and parallelism and fewer steps cut everything around the model.
The questions build in that order. The Basic questions cover where latency comes from, time to first token, tokens per second, the cost of long context and retrieval, parallel execution, and the trade-off between accuracy, latency and cost. The Advanced questions cover speeding up agents, speculative execution, measuring what users actually feel, safe concurrency for tool calls, retrieval latency, prompt caching, output length and load testing. Read it as one loop: measure each stage, attack the biggest one, and check quality did not move.
End-to-end latency is the sum of network and queueing, prefill over the input tokens, decode over every output token, and everything around the model: retrieval, tool calls, extra model calls and orchestration. Output length and the number of sequential steps usually dominate.
A single LLM call has two phases. Prefill processes all input tokens at once to build the model's internal state (the KV cache, a store of attention keys and values for every token so far). It is parallel and compute-heavy, and its time grows with input length. Decode then produces output tokens one at a time, each step depending on the previous token. Decode is limited mainly by how fast the hardware can read model weights from memory, so its time grows roughly linearly with the number of output tokens.
Around the model call sit the costs that are easy to forget. Network round trips to the provider, especially across regions. Queueing at the provider or your own server when load is high, which shows up as unpredictable spikes. Retrieval: embedding the query, searching an index and reranking results. Tool calls: a slow database or third-party API inside an agent step. Orchestration: every extra model call in a chain (classify, then retrieve, then answer, then check) adds a full prefill and decode.
The shape matters more than any single number. Sequential steps add up, while parallel steps cost only the slowest one. An agent that makes six model calls in a row, each producing 300 tokens, can easily take 20 seconds even if each call feels fast in isolation. Reasoning models add a further hidden cost: tokens spent thinking are decoded like any other output, so they delay the visible answer.
The practical consequence is that you cannot optimise latency by intuition. Trace a real request, attribute time to each stage, and attack the largest one. Teams routinely spend a week shaving 50 ms off retrieval when the answer step generates 800 tokens nobody reads.
Time to first token (TTFT) is the delay between sending a request and receiving the first streamed output token. It covers network, queueing and prefill over the whole input, and it decides how long a user stares at a spinner.
TTFT is measured from the moment your client sends the request to the moment the first content token arrives on the stream. Inside that window: the network round trip, any queueing at the provider or your gateway, and prefill, where the model processes every input token before it can predict the first output token. Prefill time grows with input length, so a 50,000-token prompt has a noticeably higher TTFT than a 2,000-token one on the same model.
It matters because humans tolerate a fast start followed by streaming far better than a long blank wait followed by a complete answer. With streaming, a response that takes 8 seconds in total can feel responsive if words appear within half a second. TTFT is the number that maps most directly onto that first impression in a chat interface.
Several things inflate it in ways teams miss. Reasoning models may spend seconds generating hidden thinking tokens before the first visible token, so the TTFT you care about is time to first visible token. Tool-calling turns may stream a tool call rather than text, so the user sees nothing until the tool runs and a second call begins. Server-side prompt caching cuts TTFT when the start of the prompt matches a recent request, and cold caches after a deploy can raise it. Measure the definition that matches the user experience and log it per request.
TTFT is not the whole story. A fast first token followed by a slow stream, or by three more agent steps, still produces a slow task. Use TTFT alongside tokens per second and total task time, and report percentiles rather than averages.
Output tokens per second depend on the model's size and architecture, the hardware's memory bandwidth, how many requests share the server, precision and quantization, and serving features such as speculative decoding. Reasoning and constrained output modes change the effective rate users see.
During decode, each new token requires reading essentially all of the model's active weights from GPU memory. That makes decode memory-bandwidth bound: a model with more active parameters, or hardware with less bandwidth, produces fewer tokens per second per request. This is why smaller models stream faster, and why mixture-of-experts models, which only activate part of their weights per token, can be faster than their total size suggests.
Load is the next factor. Servers batch many requests together so one read of the weights serves many sequences. Bigger batches raise total throughput but slow each individual stream, so the same model streams faster at 3 a.m. than at peak. Providers tune this trade-off for you; when you self-host, you tune it yourself with batch size and concurrency limits.
Precision and serving tricks help. Quantization (storing weights in fewer bits) reduces the bytes read per token and usually raises speed, at some quality risk. Speculative decoding lets a small draft model propose several tokens that the large model verifies in one pass, giving the same output distribution with fewer large-model steps when the draft guesses well. Longer contexts also slow decode slightly, because each step attends over a larger KV cache.
Finally, the rate users perceive depends on what is being generated. Hidden reasoning tokens consume decode time without appearing on screen. Constrained decoding for strict JSON usually adds little overhead, but a verbose schema forces many tokens of syntax. Vendors' advertised speeds are typically measured under favourable conditions, so benchmark with your own prompts at your expected concurrency.
Yes, mostly before the first token: prefill time grows with input length, and decode also slows slightly as the KV cache grows. Prompt caching removes much of the repeated prefill, but sending irrelevant context still costs time, money and often quality.
Prefill must process every input token before the first output token appears. In modern serving stacks prefill time grows roughly linearly with input length for typical sizes, with the attention part growing faster at very long lengths. Doubling a 20,000-token prompt to 40,000 tokens will raise TTFT noticeably, and a 200,000-token prompt can add several seconds before anything streams.
Long context also affects decode. Each new token attends over all previous tokens, so the per-token cost rises a little as context grows. The effect is smaller than on prefill, but it adds up over long outputs. On self-hosted models, large contexts also consume KV-cache memory, which reduces how many requests can run at once and so raises queueing under load.
Prompt caching changes the calculation for repeated content. If many requests share the same prefix (a long system prompt, a product manual, tool definitions), the provider or serving engine can reuse the computed state for that prefix and only prefill the new suffix. That can cut TTFT substantially for long, stable prefixes. It only works when the prefix is byte-identical, so the order of content matters.
Speed is not the only cost. Longer contexts often lower answer quality through distraction and context rot, where the model attends less reliably to facts buried in the middle. Sending a whole manual to answer a one-line question is slower, more expensive and frequently less accurate than retrieving the two relevant sections.
Yes: query embedding, search, reranking and any query rewriting add steps before generation, typically tens to hundreds of milliseconds. Done well, RAG can still reduce total latency by replacing a huge prompt with a small, relevant one.
Retrieval-augmented generation (RAG) puts a pipeline in front of the model. A typical request embeds the query (one call to an embedding model), searches a vector or hybrid index, optionally reranks candidates with a cross-encoder, and assembles the selected chunks into the prompt. Some pipelines add an LLM call before retrieval to rewrite or expand the query, which is often the most expensive stage because it is a full model call.
Each stage has its own cost profile. Embedding a short query is fast, but a network hop to a hosted embedding service adds round-trip time. Approximate nearest neighbour search is usually tens of milliseconds even over millions of vectors, though heavy metadata filters or a cold index can slow it. Reranking cost grows with the number of candidates and their length, so reranking 100 long chunks can take longer than the search itself.
The trade is that retrieval shrinks the prompt. Sending 4,000 relevant tokens instead of 80,000 raw tokens cuts prefill time, cost and distraction. For many applications the net effect is a faster and better answer. RAG makes things slower overall mainly when the corpus would have fit cheaply in a cached prompt anyway, or when the pipeline grows extra LLM stages that each add a second or more.
Retrieval latency also shapes perceived speed. The user sees nothing until retrieval finishes and generation starts, so retrieval time adds directly to TTFT. Show sources or a status line as soon as retrieval returns, and keep retrieval steps that can run in parallel (keyword and vector search, for example) concurrent.
Parallel execution runs independent operations at the same time, so the step costs as long as the slowest operation instead of the sum of all of them. In LLM systems it applies to retrieval sources, tool calls and independent model calls.
Most LLM work is I/O bound: waiting on a provider API, a database or a search service. While one request waits, the program can issue others. With concurrency (for example asyncio.gather in Python or Promise.all in TypeScript), three 400 ms lookups finish in about 400 ms rather than 1,200 ms. Strictly, this is concurrency rather than CPU parallelism, but the latency effect is the same and teams call it parallel execution.
The rule for what can run together is dependency. If step B needs the output of step A, they must run in order. If they only share an input, they can run concurrently. Fetching a customer record and their open orders by customer ID can run together. Looking up the order and then the shipment for that order cannot. Drawing the dependency graph of a request usually reveals two or three sequential steps that never needed to be sequential.
Parallelism applies at several levels. Within retrieval: keyword and vector search together. Within an agent: several read-only tool calls issued in one turn, which many provider APIs support. Across model calls: generating three independent sections of a report, or running a classifier and a safety check at the same time as retrieval. This last pattern is often called fan-out/fan-in.
It has costs. Concurrent calls consume rate limits faster, raise peak load on downstream services, and complicate error handling, because one failure among five needs a policy (retry it, drop it, or fail the whole step). Add per-call timeouts and a concurrency cap so a burst of parallel work does not trip a provider's rate limit and make everything slower.
Stronger models, more reasoning, extra retrieval and verification steps usually raise accuracy but add time and cost. The skill is spending that budget only on requests where the extra accuracy changes the outcome.
Almost every quality lever in an LLM system costs time and money. A larger model is slower per token and more expensive. Higher reasoning effort generates more hidden tokens. Retrieving more documents enlarges the prompt. A verification pass or a reviewer model adds a full extra call. Sampling several answers and voting multiplies cost. Each can improve accuracy, but none is free.
The trade-off is not linear. Accuracy gains often flatten after the first improvement, while cost and latency keep rising. Going from no retrieval to good retrieval may lift accuracy a lot; doubling the number of retrieved chunks after that may add little. Measuring the curve on your own eval set, rather than assuming bigger is better, shows where the knee is.
The way out is to stop treating all requests the same. Routing sends simple requests to a fast model and hard ones to a strong model. A cascade tries the fast path first and escalates only when a confidence check or validator fails. Tiered review applies an expensive check only to high-stakes outputs, such as refunds over a threshold or medical content. The average request becomes fast and cheap while the important ones still get the full treatment.
Users also weigh this differently by task. A coding agent fixing a bug can take a minute if it gets it right; an autocomplete suggestion must appear in a few hundred milliseconds or it is useless. Write the latency target and accuracy target together in the product requirements so the trade-off is a decision, not an accident.
Profile the trace, then cut sequential model calls: remove unnecessary steps, batch independent tool calls in one turn, cache safe reads, keep contexts small, use smaller models for routine steps and set clear stop rules. Faster tokens help less than fewer turns.
An agent's latency is roughly the number of loop iterations times the cost of each iteration. Each iteration is a full model call (prefill over a growing context, then decode) plus whatever tools it triggers. A 12-step agent at 2.5 seconds per step takes 30 seconds before any tool time. So the biggest gains usually come from fewer iterations, not faster ones.
Start by reading traces of slow runs. Common waste: the agent calls a search tool, reads the result, then calls it again with a slightly different query; it re-reads a file it already has; it spends a turn announcing a plan; it keeps going after the task is done because the stop condition is vague. Fixes include better tool design (a tool that returns the needed fields in one call instead of three), instructions that encourage issuing all independent reads in one turn, and an explicit definition of done.
Next, make each iteration cheaper. Keep the context small by summarising or truncating old tool results, since every turn re-prefills the entire history. Keep the stable prefix (system prompt, tool definitions) identical across turns so prompt caching applies. Use parallel tool calls, which most provider APIs support, so several reads in one turn cost one round trip. Route routine steps such as classification, extraction or formatting to a smaller, faster model, and keep the strong model for planning and hard decisions.
Then remove the agent where it is not needed. If the trace shows the same sequence of steps every time, that part is a workflow, and code can run it deterministically without model calls between steps. Many fast agents are a workflow with a model at one or two decision points. Finally, cache safe work: read-only tool results with a short time-to-live, and repeated lookups within one session.
Speculative execution starts work that will probably be needed before you know for sure, such as retrieving documents while intent classification is still running, and discards it if the guess was wrong. It trades some wasted compute for lower latency.
The idea comes from CPU design: rather than wait for a branch to resolve, start on the likely path and throw the work away if the prediction fails. In an LLM pipeline the expensive dependency is usually a model call whose output decides what happens next. Instead of waiting for it, you start the most likely follow-up immediately and keep the result only if the decision agrees.
Common patterns: run retrieval in parallel with an intent classifier, because most requests will need retrieval anyway. Start a fast-model answer while a router decides whether the request needs the strong model, and use the fast answer if the router says it is sufficient. Prefetch the likely next page or tool result while the user reads the current answer. Begin generating a draft while a guardrail check runs, and only release the stream once the check passes. In voice agents, start forming a response on a likely end of turn and cancel it if the user keeps talking.
Do not confuse this with speculative decoding, which happens inside the model server: a small draft model proposes several tokens and the large model verifies them in one pass, preserving the output distribution. Speculative execution is an application-level pattern you build in your orchestration code.
The cost is waste. Every wrong guess burns tokens and downstream capacity, so it pays when the guess is right most of the time, the speculative work is cheap relative to the latency saved, and it has no side effects. Never speculate on writes such as sending emails, creating tickets or charging cards. Track the hit rate: if retrieval is wasted on 40% of requests, the latency gain may not justify the extra load. Cancel losing work promptly so it stops consuming rate limits.
Measure what the user experiences: time to first useful content, time until the task is complete, and how often they wait without feedback. Instrument from the client, report percentiles, and pair speed metrics with abandonment and retry rates.
Model-side metrics (TTFT, tokens per second) describe the server, not the user. A user cares when something useful appears and when they can act. Time to first useful content is when the first meaningful text, source or partial result is on screen, which may be later than the first token if the model opens with filler. Time to task completion is when the user has what they came for, which for an agent may be after several steps and a confirmation.
Instrument from the client, not only the server, because client-side numbers include network, rendering and any front-end waiting. Attach a request ID from the browser or app to the server trace so you can line up what the user saw with where the time went. Record a few timestamps per interaction: request sent, first content rendered, final content rendered, and the user's next action.
Report percentiles. p50 tells you the typical experience; p95 and p99 tell you about the users who will complain. LLM latency has a long tail, driven by long outputs, retries, queueing and agent loops, so an average can look fine while one user in twenty waits 30 seconds. Segment by request type as well: a mixed dashboard hides the fact that one workflow is slow.
Then connect speed to behaviour. Abandonment (closing or navigating away before completion), regenerate or retry clicks, and stop-generation clicks reveal whether the latency hurts. Perceived speed also depends on feedback: a progress message naming the current step, sources shown as soon as retrieval returns, and streaming all make the same wall-clock time feel shorter. Test those changes like any product change, comparing abandonment before and after.
Run tool calls concurrently when they are independent: neither needs the other's output, they do not modify shared state, and their combined load stays within rate limits. Reads usually qualify; writes, and anything order-dependent, should run in sequence.
Many provider APIs let a model request several tool calls in one turn. Your harness decides whether to execute them in parallel. The default should be concurrency for calls that are independent and read-only: fetching weather and venue availability, searching two indexes, reading three files. These finish in the time of the slowest call and return to the model in one round trip.
Sequence calls when there is a data dependency (the second call needs an ID from the first), a state dependency (both write to the same record, or one reads what the other writes), or an ordering requirement (reserve then charge, never the reverse). Running dependent writes concurrently causes race conditions such as double bookings, lost updates and charges without reservations, and these rarely appear in tests because test runs are slow and sequential.
A practical policy tags each tool with metadata: whether it is read-only, whether it is idempotent (safe to repeat), and which resource it touches. The harness then runs read-only calls concurrently, runs writes one at a time or serialised per resource, and refuses to parallelise anything without the tags. This keeps the decision in code rather than trusting the model to reason about concurrency.
Concurrency also needs limits. A model may emit ten calls at once; executing them all could exceed a downstream rate limit or overload a small internal service. Use a semaphore or pool to cap concurrency, set per-call timeouts, and decide how partial failures are reported back to the model, typically as structured error results for the failed calls alongside the successful ones.
Time each stage, then fix the slowest: co-locate services, filter early, tune the index for your recall target, run independent searches concurrently, rerank a small candidate set, and cache embeddings and frequent queries. Always check recall did not drop.
Retrieval is a pipeline, so start with per-stage timing: query rewriting, embedding, search, filtering, reranking and chunk fetching. Often one stage dominates. A hosted embedding call in another region, a reranker scoring 200 candidates, or an LLM query-rewrite step can each cost more than the vector search itself.
Index and search. Approximate nearest neighbour indexes such as HNSW (hierarchical navigable small world graphs) trade recall for speed through parameters like the number of candidates explored per query. Raising that parameter improves recall and slows search, so tune it against a labelled recall set rather than by feel. Apply metadata filters (tenant, language, date) in a way your database can use efficiently; some engines filter during graph traversal, while others filter afterwards and may return too few results or scan too much. Keep the index in memory where feasible, and smaller embedding dimensions or quantized vectors reduce memory and search time at some recall cost.
Reranking is usually the biggest lever. A cross-encoder reranker scores each query-chunk pair, so its cost scales with candidate count and chunk length. Retrieving 50 and reranking to the top 8 is a common shape; if recall at 30 is nearly the same as at 100, rerank 30. Truncate long chunks for scoring, run the reranker close to the application, and batch the pairs in one call.
Caching and concurrency. Cache query embeddings for repeated or popular queries, cache final results for identical normalised queries with a short time-to-live, and run keyword and vector search concurrently in hybrid setups. Skip retrieval entirely when a cheap router says the question does not need it. After each change, rerun the retrieval eval (recall at k, answer quality) so a faster pipeline is not a worse one.
Prompt caching reuses the computed attention state for a prompt prefix that matches a recent request, so the model only prefills the new tokens. It can cut time to first token sharply for long, stable prefixes, but only if the prefix is identical byte for byte.
During prefill the model computes keys and values for every input token and stores them in the KV cache. If the next request starts with exactly the same tokens, that work does not need to be repeated. Prompt caching (also called prefix caching) stores the KV state of a prefix and reuses it, so a request with a 30,000-token shared prefix and a 500-token new suffix pays prefill on roughly 500 tokens. Both hosted providers and self-hosted serving engines offer forms of it.
The matching rule is strict: the cache applies to the longest identical prefix, and the first differing token ends reuse for everything after it. That makes prompt layout the main design decision. Put the most stable content first (system instructions, tool definitions, reference documents), then semi-stable content (conversation history), then the newest user message last. A timestamp, request ID or per-user greeting near the top invalidates the cache for every request.
Providers differ in the mechanics, so check their documentation. Some cache automatically above a minimum prompt length; others require you to mark cache breakpoints explicitly. Caches expire after a period of inactivity, typically minutes, and some offer longer retention at different pricing. Cached input tokens are usually billed at a discount, so the same change often improves both latency and cost. Responses typically report how many input tokens were read from cache, which you should log.
Prompt caching is different from response caching or semantic caching, which reuse a whole previous answer for a matching question. Prompt caching still runs the model and produces a fresh answer; it only skips repeated prefill. It is safe to use for personalised responses, and it matters most for agents and long chats, where every turn re-sends a growing history behind a large fixed prefix.
Replay realistic prompts at increasing concurrency, measure TTFT, tokens per second, end-to-end latency and error rates at each level, and find where p95 latency or rate-limit errors break your targets. Synthetic one-line prompts give misleading results.
LLM services behave differently under load from ordinary web services. Throughput is limited by tokens, not requests: one request with a 50,000-token prompt and a 2,000-token answer can cost as much server time as hundreds of short ones. Providers enforce limits on requests per minute and tokens per minute, and self-hosted servers slow each stream as batch sizes grow. A load test must therefore use a realistic distribution of input and output lengths, ideally sampled from production logs with sensitive data removed.
Ramp concurrency in steps (for example 1, 5, 20, 50, 100 simultaneous sessions) and hold each step long enough to reach a steady state. At each step record TTFT, output tokens per second per stream, end-to-end latency at p50, p95 and p99, error rates by type (rate limit, timeout, server error), and for self-hosted models, GPU memory and queue depth. The useful output is a curve showing where latency starts to climb steeply, and which limit causes it.
Test the whole application, not only the model. Retrieval services, databases, tool APIs and your gateway often saturate before the model does. Agent workloads need session-level tests, since one user action can produce many model and tool calls. Include streaming: many load tools measure time to the complete response by default, which hides TTFT behaviour.
Be careful with hosted providers. Load testing a production API consumes your real quota and may affect production traffic sharing the same limits, so use a separate project or key where the provider allows it and agree large tests in advance if required. Results from a hosted API also vary by time of day as provider load changes, so repeat key runs.
Generate fewer tokens: ask for concise formats, cap output length, trim verbose structured output, lower reasoning effort where it does not help, and split long outputs into parts that can stream or run in parallel. Output length is usually the largest latency lever.
Each output token is a sequential decode step, while input tokens are processed in parallel during prefill. As a rough rule, generating one extra output token costs far more time than reading one extra input token. An answer of 800 tokens at a typical streaming speed takes several times longer than one of 200 tokens, whatever happens before it. That makes output length the first thing to examine when an answer feels slow.
Most verbosity is optional. Instruct the model on the format: 'three bullet points', 'one sentence then the command', 'no preamble or summary'. Show a short example of the desired output, since models copy examples more reliably than they follow length instructions. Set a maximum output token limit as a safety net, but do not rely on it for brevity; a hard cutoff truncates mid-sentence rather than producing a shorter answer.
Structured output has its own overhead. Long key names, deep nesting and fields the application never reads all cost decode steps. Ask only for the fields you use, and prefer short keys for high-volume internal calls. For reasoning models, the hidden thinking tokens are output too: lowering reasoning effort on simple tasks can cut seconds without changing the visible answer quality, which you should confirm on your eval set.
When a long output is genuinely needed, change its shape. Stream it so reading starts early. Generate independent sections in parallel and assemble them. Return a short answer first with an option to expand. Each of these keeps the total work similar but reduces how long the user waits for what they need.