Understand the compute and operating costs behind running your own models.
Self-hosting means your team, not an API provider, runs the model: you pick the weights, load them onto GPUs you rent or own, put a serving engine in front, and operate the endpoint like any other production service. Open-weight models made this practical, because their trained parameters can be downloaded and run anywhere their license allows. The appeal is control over data, versions, latency and customization. The cost is that every problem a provider used to absorb (capacity planning, batching, failover, upgrades, utilization) is now yours.
The mental model that explains almost everything in this chapter is that LLM inference is a memory problem before it is a compute problem. The weights must sit in GPU memory (VRAM), every active conversation needs its own KV cache next to them, and generating each token means streaming the weights through the chip. Quantization shrinks the weights, continuous batching and paged KV caches let many requests share one copy of them, and tensor or pipeline parallelism spread a model that is too big for one GPU across several. How many users one GPU can serve, and what a token really costs you, both fall out of that memory arithmetic.
The questions build in that order. The basic group defines self-hosting, why teams do it, what VRAM is and why it is the binding constraint, how quantization trades precision for memory, what a serving stack does, and what the KV cache is. The advanced group covers the techniques that make serving efficient (continuous batching, tensor and pipeline parallelism), how to size capacity, when a managed API is the better choice, and how to compare costs honestly. Two added questions cover the licensing check that comes before any of this and the operational work that comes after it.
Self-hosting a model means running inference on compute you control, such as your own GPUs, cloud GPU instances or a private cluster, using downloaded open weights and a serving engine, instead of calling a provider's hosted API.
With a hosted API, you send a request and a provider handles everything behind it: the weights, the GPUs, the batching, the scaling and the uptime. When you self-host, you take over that stack. You download a model's weights (the billions of trained numbers that define it), load them into GPU memory, run a serving engine that turns requests into batched GPU work, and expose an HTTP endpoint that your applications call. Many open-source serving engines expose an OpenAI-style chat completions interface, so client code often changes little.
Self-hosting is a spectrum, not a single choice. At one end you rent a GPU virtual machine and run an engine on it. In the middle you deploy containers onto a Kubernetes cluster with GPU nodes, or use a managed platform that runs open weights inside your cloud account. At the other end you buy servers for your own data centre or run on-premises in an air-gapped network. What they share is that you decide which exact model version runs, where the data goes, and how capacity is provisioned.
Only open-weight models can be self-hosted. Their parameters are published for download under a license, though the training data and code usually are not. The most capable proprietary models are typically available only through their providers' APIs or cloud partners, so choosing to self-host also narrows the set of models you can use. That trade-off is the first thing to check against your quality bar.
The work does not stop at a running endpoint. A production deployment needs health checks, autoscaling, monitoring of latency and GPU utilization, a way to roll out a new model version safely, and redundancy so one failed GPU does not take the feature down. Those are the costs teams most often underestimate.
Companies self-host for data control, regulatory or residency requirements, version stability, deep customization such as fine-tuned or specialized models, predictable latency, and lower unit cost when traffic is high and steady enough to keep GPUs busy.
The most common driver is data control. Some organizations cannot send certain data to an external processor at all, whether because of contracts, sector rules, data residency requirements (keeping data inside a country or region) or an air-gapped environment. Enterprise API agreements with no-training and zero-retention terms solve this for many teams, so check whether that is enough before taking on a GPU fleet.
The second driver is control over the model itself. A hosted model can be updated or retired on the provider's schedule, which can shift behaviour your prompts and evals depend on. Self-hosting pins the exact weights until you choose to change them. It also allows customization that APIs may restrict: full fine-tuning, many LoRA adapters served on one base model, custom decoding or constrained generation, and direct access to log probabilities and internal states.
The third driver is economics at sustained scale. An API charges per token, and that price includes the provider's margin and idle capacity. A GPU you rent costs the same per hour whether it is busy or idle. If your traffic is large and steady, and the job can be done well by a smaller model, a well-utilized GPU can produce tokens more cheaply than an API. At low or spiky volume the opposite is usually true.
Latency and availability can also favour self-hosting: a model in the same region or even on the same device avoids network hops and shared-tenant rate limits. The counterweights are real. You need people who can operate GPU infrastructure, you are limited to open-weight models, and you carry the cost of redundancy and idle capacity yourself.
VRAM is the high-bandwidth memory attached to a GPU. For LLM inference it holds the model weights, the KV cache for every active request, and temporary working buffers, and both its size and its bandwidth limit what you can serve.
A GPU has its own memory, separate from the server's system RAM. On data-centre accelerators this is usually HBM (high-bandwidth memory) stacked next to the chip; on consumer and workstation cards it is GDDR memory. Capacities range from a few gigabytes on laptop GPUs, through 16 to 32 GB on consumer and workstation cards, to 80 GB and beyond on data-centre parts. The exact numbers change with each hardware generation, so check current specifications.
For inference, three things compete for VRAM. The weights take a fixed amount: roughly the parameter count times the bytes per parameter, so an 8-billion-parameter model at 16-bit precision needs about 16 GB. The KV cache grows with the number of active requests and their context lengths. Activations and workspace (temporary tensors, the CUDA context, the serving engine's buffers) take a few more gigabytes. Serving engines usually pre-allocate a configured fraction of VRAM at startup and divide what remains after the weights into KV cache blocks.
Capacity is only half the story. Memory bandwidth, measured in terabytes per second on modern data-centre GPUs, decides how fast tokens come out. Generating one token for one request requires reading essentially all the weights from VRAM once, so a single stream's decode speed is roughly bandwidth divided by weight size. That is why a GPU with more bandwidth can be faster for LLMs even when its raw compute figure looks similar.
System RAM and CPU offloading exist, but moving data across the PCIe bus is an order of magnitude slower than reading HBM, so offloaded layers make generation dramatically slower. Treat VRAM as the real budget.
VRAM decides whether a model runs on a GPU at all, how many requests it can serve at once, and how long their contexts can be. When it runs out, you must quantize, shard across GPUs, shorten contexts, or accept slower offloading.
The first question VRAM answers is fit. Weight memory is roughly parameters times bytes per parameter: 2 bytes at 16-bit, 1 byte at 8-bit, about half a byte at 4-bit. A 70-billion-parameter model therefore needs about 140 GB at 16-bit, which exceeds any single common GPU, but roughly 35 to 40 GB at 4-bit, which fits on one 48 GB or 80 GB card. That arithmetic is often what decides between one GPU and a multi-GPU setup.
The second question is concurrency. Once the weights are loaded, the rest of VRAM holds KV caches, and every active request needs its own. If each request uses 2 GB of KV cache and you have 40 GB free, you can run about 20 requests at once; the 21st waits in a queue. More free memory means bigger batches, and bigger batches are how a GPU reaches good throughput, because the cost of reading the weights is shared across every request in the batch.
The third is context length. KV cache grows linearly with tokens in context, so a deployment that comfortably serves 50 users at 4,000 tokens may only serve a handful at 100,000 tokens. Long-context features are a memory decision, not just a model capability.
When memory runs short, the options all trade something away: quantize weights or the KV cache (some quality risk), split the model across GPUs with tensor or pipeline parallelism (more hardware and interconnect cost), cap context length or concurrency (product limits), or offload to CPU memory (much slower). Knowing which constraint you are hitting tells you which lever to pull.
Quantization stores model weights, and sometimes activations or the KV cache, in fewer bits, such as 8 or 4 instead of 16. It cuts memory and often speeds up generation, at a quality cost that is usually small at 8-bit and must be measured at 4-bit and below.
Models are typically trained and released in 16-bit floating point (FP16 or BF16). Quantization maps those values onto a smaller set of numbers: 8-bit integers or 8-bit floats (FP8), or 4-bit formats, with a scale factor per group of weights so that the small integers can be converted back to approximately the original values during computation. Halving the bits halves weight memory, and because decoding is limited by how fast weights are read from memory, smaller weights also mean faster tokens on bandwidth-bound workloads.
There are several flavours. Weight-only quantization (methods such as GPTQ and AWQ, and the quantized files used by local runtimes such as GGUF) compresses weights and dequantizes them on the fly; it mainly saves memory and helps decode speed. Weight-and-activation quantization, such as INT8 or FP8 matrix multiplication on hardware that supports it, also speeds up compute-heavy prefill. KV cache quantization stores keys and values in 8 bits or fewer to fit more concurrent context. Most are applied after training (post-training quantization) using a small calibration set, so no retraining is needed.
The quality cost is real but uneven. 8-bit is close to lossless for most models and tasks. 4-bit usually loses a little on general benchmarks but can lose more on precise tasks such as maths, code, structured output, or less common languages. Below 4 bits, degradation grows quickly. Larger models tolerate quantization better than small ones, which is why a 4-bit large model often beats a 16-bit model of half the size in the same memory.
Results also vary by method and by implementation, so a published benchmark for one quantized file says little about another. Evaluate the exact artifact you will deploy on your own task.
Inference serving is running a model as a production endpoint: a serving engine schedules and batches requests onto GPUs, manages KV cache memory, streams tokens back, and the surrounding platform handles routing, scaling, auth and monitoring.
Loading a model in a notebook and calling generate handles one request at a time and wastes most of the GPU. A serving engine is the software built to do this efficiently for many users. Open-source engines in this space (vLLM, SGLang, Hugging Face Text Generation Inference, NVIDIA's TensorRT-LLM with Triton, and llama.cpp for CPU and edge devices, among others) differ in features and hardware support, but the core jobs are the same.
Inside the engine, a scheduler decides which requests run on each step, using continuous batching so new requests join without waiting for others to finish. A KV cache manager allocates memory for each request's context, often in fixed-size pages to avoid fragmentation, and can reuse cached prefixes shared across requests, such as a long system prompt. Optimized kernels perform attention and matrix multiplications, including quantized formats. The engine streams tokens as they are generated and enforces limits such as maximum context and output length.
Around the engine sits the platform. A gateway handles authentication, rate limiting and request size caps. A router or load balancer spreads traffic across replicas, ideally aware of queue depth and cached prefixes rather than simple round-robin. Autoscaling adds replicas when queues grow, bearing in mind that a cold start means pulling tens of gigabytes of weights and warming up, which can take minutes. Monitoring tracks time to first token, tokens per second, queue time, GPU memory and error rates.
The key trade-off a serving setup tunes is throughput versus latency. Bigger batches use the GPU better and lower cost per token, but each request gets a smaller share of each step, so per-user speed drops. Most teams set a latency target and push batch size as high as that target allows.
The KV cache stores the attention keys and values already computed for earlier tokens, so each new token only computes its own instead of reprocessing the whole sequence. It makes generation fast but consumes VRAM in proportion to context length and concurrency.
In a transformer, each token at each layer produces a query, a key and a value vector. To generate the next token, its query is compared with the keys of all earlier tokens, and the result weights their values. Without a cache, generating token 1,001 would mean recomputing keys and values for the first 1,000 tokens at every layer, and doing so again for token 1,002. The KV cache stores them once, so each step only computes the new token's vectors and reads the rest from memory.
This splits inference into two phases. Prefill processes the whole prompt in parallel and writes its keys and values into the cache; it is compute-heavy and determines time to first token. Decode generates one token at a time, reading the weights and the growing cache at each step; it is memory-bandwidth-bound and determines tokens per second.
The cost is memory. Per token, the cache holds 2 (key and value) times layers times KV heads times head dimension times bytes. For a typical 8B model with grouped-query attention that is about 128 KB per token at 16-bit, so a 32,000-token conversation needs about 4 GB. Older designs without grouped-query attention can need four to eight times more. Multiply by concurrent users and the cache can exceed the weights.
Serving engines manage this carefully. Paged attention, introduced with vLLM in 2023, stores the cache in fixed-size blocks like virtual memory pages, cutting waste from fragmentation and over-reservation. Prefix caching shares the blocks for a common prompt prefix across requests, so a long system prompt is computed once. KV cache quantization and eviction or offloading policies stretch it further. Provider-side prompt caching discounts are an API-level view of the same idea.
Continuous batching schedules work one generation step at a time, so finished requests leave the batch and waiting requests join immediately instead of waiting for the whole batch to finish. It keeps the GPU full and can raise throughput several-fold over static batching.
Decoding reads the full set of weights from memory to produce one token. If one request is running, that expensive read serves one token; if 64 requests are batched, the same read serves 64 tokens. Batching is therefore the main lever for throughput. The difficulty is that LLM requests have wildly different lengths: one answer finishes after 20 tokens, another runs to 800.
With static batching, the engine collects a batch, runs it until the longest request finishes, then starts the next batch. Short requests sit finished but occupying a slot, new arrivals wait, and the GPU spends much of its time computing for a handful of stragglers. Continuous batching (also called iteration-level or in-flight batching, described in the 2022 Orca paper) makes the scheduling decision at every decode step. When a request emits its end token, its slot and KV cache memory are released; on the next step a queued request is admitted. The batch stays full as long as there is demand and memory.
Admitting a new request means running its prefill, which is compute-heavy and can stall the decode steps of everyone else, causing visible stutter in streamed output. Modern engines use chunked prefill, splitting a long prompt into pieces processed alongside ongoing decodes, to smooth this. Some deployments go further with prefill-decode disaggregation, running the two phases on separate GPU pools and transferring the KV cache between them, so each can be tuned for its own bottleneck.
Continuous batching works hand in hand with paged KV cache management, because admitting and evicting requests every step needs memory that can be allocated and freed in small blocks. When memory runs out mid-generation, engines either queue new arrivals or preempt a running request and recompute it later, which shows up as latency spikes.
The tuning knobs are the maximum number of concurrent sequences and the maximum tokens processed per step. Raising them increases throughput until per-request latency breaches your target. The right values come from load testing with your real length distribution.
Tensor parallelism splits the matrices inside each layer across several GPUs, so each GPU holds a slice of every layer and computes part of every token, combining results with fast collective communication. It lets a model exceed one GPU's memory and can cut per-token latency.
A transformer layer is mostly large matrix multiplications. Tensor parallelism, popularized by the Megatron-LM work, partitions those matrices. In the attention block, different GPUs take different heads. In the feed-forward block, the first matrix is split by columns and the second by rows, so each GPU computes a partial result on its own slice. After each block, the GPUs exchange and sum their partial results with an all-reduce operation before moving on.
Because every GPU holds a fraction of every layer, weight memory per GPU drops roughly in proportion to the parallel degree. A 70B model at 16-bit (about 140 GB) split four ways needs about 35 GB of weights per GPU, leaving room on 80 GB cards for a large KV cache, which is also sharded by head. And because each GPU reads only its slice of the weights per token, decoding can be faster than on a single, larger device, since the aggregate memory bandwidth is higher.
The cost is communication. There are typically two all-reduces per layer per step, so a model with 80 layers performs 160 synchronizations for every token generated. This only works well over a very fast interconnect such as NVLink or similar GPU-to-GPU links inside one server. Over ordinary PCIe, and especially across servers over a network, communication time can dominate and throughput falls sharply. That is why tensor parallelism is usually confined to the GPUs within a single node, with degrees of 2, 4 or 8.
There are also structural constraints. The parallel degree usually needs to divide the number of attention heads (and of KV heads under grouped-query attention), and returns diminish: going from 4 to 8 GPUs adds communication overhead while each GPU does less useful work. Mixture-of-experts models add expert parallelism, placing different experts on different GPUs, which has its own all-to-all communication pattern.
Use tensor parallelism when a model does not fit on one GPU, or when you need lower per-token latency than one GPU can provide. If the model fits comfortably on one GPU, running several independent single-GPU replicas usually gives better total throughput.
Pipeline parallelism assigns consecutive groups of layers to different GPUs or nodes, passing activations from one stage to the next. It needs far less communication than tensor parallelism, so it suits slower links between servers, but adds per-token latency and idle bubbles.
Instead of splitting every layer, pipeline parallelism splits the model by depth. With an 80-layer model on two GPUs, GPU 1 runs layers 1 to 40 and GPU 2 runs layers 41 to 80. After GPU 1 finishes its layers for a batch, it sends the resulting activations (one vector per token, small compared with the weights) to GPU 2. Communication happens once per stage boundary rather than twice per layer, which is why pipeline parallelism tolerates slower interconnects, including networking between servers.
The catch is sequential dependency. If one batch flows through the pipeline, GPU 2 sits idle while GPU 1 works, and vice versa: these idle periods are called pipeline bubbles. Engines reduce them by splitting work into micro-batches so that while GPU 2 processes micro-batch A, GPU 1 is already working on micro-batch B. With enough concurrent requests, all stages stay busy and throughput scales reasonably well.
Latency does not improve. A single token still has to traverse every stage in sequence, and each hop adds transfer time, so per-token latency is roughly the same as one big GPU plus communication overhead, and somewhat worse in practice. Each GPU still reads only its own layers' weights per token, so memory per GPU drops in proportion to the number of stages.
In practice the two techniques are combined. The common pattern for very large models is tensor parallelism within each server, where GPUs share a fast interconnect, and pipeline parallelism across servers, where the links are slower. For example, a model too large for one 8-GPU node might run as tensor-parallel 8 within each of two nodes, with a pipeline degree of 2 between them.
Pipeline stages must also be balanced. Layers are usually uniform, but the first and last stages carry the embedding and output layers, and an unbalanced split leaves one GPU as the bottleneck while others wait.
It depends on model size, context and output lengths, latency targets and how often users actually send requests. Estimate concurrency from free KV cache memory and Little's law, then confirm with a load test using real traffic shapes.
"Users" is the wrong unit for a GPU. What a GPU handles is concurrent active requests, and how many of those it can hold depends on two limits. The memory limit is how many KV caches fit: free VRAM after weights, divided by cache per request at your typical context length. The latency limit is how many requests can share each decode step before per-request tokens per second fall below what your product needs. Whichever limit you hit first is your ceiling.
Converting concurrent requests into users needs Little's law: average concurrency equals arrival rate times average time in the system. If a typical request takes 8 seconds end to end and the GPU can hold 40 at once while meeting your latency target, it can sustain about 5 requests per second. If an active user sends one message per minute at peak, that is about 300 simultaneously active users, which might correspond to thousands of registered users depending on how many are online at peak.
The traffic shape matters enormously. A classifier with 300 input tokens and a 5-token output finishes in a fraction of a second and can serve hundreds of requests per second. A chat assistant with 6,000-token contexts and 500-token answers holds memory for many seconds per request. An agent that makes ten model calls per task multiplies load by ten. Long outputs are expensive because decode is sequential; long inputs are expensive because of prefill compute and KV memory.
Arithmetic gets you within a factor of two; a load test gets you the real number. Replay a sample of real prompts with their real length distribution, ramp concurrency in steps, and record time to first token, tokens per second and error rate at the p50 and p95. The supported load is the highest step where the p95 still meets your target. Repeat after any change to model, quantization, engine version or prompt template, because each moves the number.
Then plan for peaks and failure. Size for the busiest hour, not the daily average, and keep enough spare capacity that losing one replica does not push the rest past their limit.
An API is usually better when traffic is low, spiky or unpredictable, when you need the strongest proprietary models, when you lack GPU operations skills, or when time to market matters more than unit cost. You pay per token and never pay for idle hardware.
The core economic difference is that an API is variable cost and a GPU is fixed cost. A provider pools demand from many customers, so their GPUs stay busy and you pay only for the tokens you use. A self-hosted GPU costs the same per hour whether it serves a thousand requests or none. If your traffic is concentrated in business hours, varies week to week, or is still small, most of the GPU-hours you pay for will be idle, and the effective cost per token can easily be several times the API price.
Quality is the second reason. The most capable frontier models are generally available only through their providers. Open-weight models have closed much of the gap and are strong for many tasks, but for hard reasoning, long-horizon agents or the newest capabilities, the best API model may still be measurably better on your eval set. A cheaper model that fails more tasks is not cheaper per successful task.
Operations is the third. Running inference reliably means GPU capacity planning, driver and engine upgrades, redundancy across zones, autoscaling with multi-minute cold starts, security patching and on-call coverage. An API provider absorbs all of that, along with features such as prompt caching, batch discounts, structured output and tool calling that you would otherwise implement yourself. For a small team, engineer time is often the largest hidden cost of self-hosting.
Data concerns, which often push teams towards self-hosting, are frequently addressed by enterprise API terms: commitments not to train on your data, zero or limited retention, regional processing, and the same models offered through major cloud platforms inside your existing cloud agreement. Check those options before assuming self-hosting is the only compliant path.
The common mature pattern is hybrid: an API for low-volume, high-difficulty and experimental work, and self-hosted smaller models for high-volume, well-defined tasks once traffic is proven. Keeping both behind one internal interface makes it easy to move a workload when the numbers change.
Compare cost per successful task, not GPU-hour price against token price. Include GPUs at realistic utilization, redundancy and idle hours, engineering and on-call time, storage, networking and evaluation, and adjust for any quality difference between the models.
The tempting comparison divides a GPU's hourly price by its peak throughput and sets that against an API's per-token price. That number is almost always too optimistic, for three reasons. Peak throughput was measured with short, uniform requests at maximum batch size, not your traffic. GPUs are not busy every hour, and utilization is the single largest multiplier: a GPU busy 25% of the time costs four times as much per token as the same GPU at full load. And production needs redundancy: at least two replicas, often in two zones, plus headroom for peaks.
A fair comparison works from your workload upward. Start with monthly volume as tasks, not tokens, and the token profile per task. Measure throughput of the candidate model on your hardware with a load test at your latency target. From that, compute the replicas needed at peak, add redundancy, and multiply by hours actually provisioned (reserved around the clock, or scaled by schedule). Then add the non-GPU costs: storage for weights and logs, networking, the gateway and monitoring stack, and engineering time for setup, upgrades and on-call, which for a small deployment can rival the GPU bill.
On the API side, use your measured token counts, including retries and agent loops, with prompt caching and batch discounts applied where they genuinely fit your workload. Check the provider's current pricing rather than relying on remembered figures, since prices change often.
Then normalize by quality. If the self-hosted model passes 90% of tasks on your eval set and the API model passes 96%, the failures have a cost too: retries, human review, escalation or lost users. Dividing total cost by successful tasks puts both options on the same footing. The break-even point usually appears at high, steady volume on tasks where a smaller open model reaches the quality bar, often after fine-tuning.
Finally, run a sensitivity check. Vary utilization, volume and quality by plausible amounts and see whether the decision flips. A decision that holds across reasonable ranges is safe; one that depends on hitting 80% utilization is a bet.
Check whether commercial use is allowed, any user-count or revenue thresholds, acceptable-use restrictions, rules on using outputs to train other models, attribution and naming requirements, and what obligations pass to fine-tuned derivatives you distribute.
"Open" covers a wide range of terms. Some model weights are released under standard permissive software licenses such as Apache 2.0 or MIT, which allow commercial use, modification and redistribution with light conditions. Many others ship under custom model licenses written by the releasing organization. These can be generous in practice but carry conditions that standard open-source licenses do not, and they differ from one model family to the next, sometimes between versions of the same family.
The clauses that matter most in practice are: whether commercial use is permitted at all (some releases are research or non-commercial only); scale thresholds, where organizations above a stated number of users or revenue need a separate agreement; an acceptable use policy listing prohibited applications, which becomes a contractual obligation on you; restrictions on using outputs to train or improve other models, which matters if you plan distillation; attribution and naming rules, such as including the original name in a derivative's name; and jurisdiction or field-of-use limits.
Derivatives need particular care. Fine-tuned weights, merged models and quantized files usually inherit the base model's license, and if you redistribute them, including inside a product shipped to customers' devices, the obligations travel with them. Community-uploaded quantized or fine-tuned versions on model hubs may not state their lineage accurately, so trace each one back to the original release.
Terminology is also contested. The Open Source Initiative published an Open Source AI Definition in 2024 that expects, among other things, sufficient information about training data, which most open-weight releases do not provide. Calling a model "open source" in contracts or marketing can therefore be inaccurate; "open-weight" is the safer term. None of this is legal advice: have counsel review the license for any model that goes into a commercial product.
Beyond a running endpoint you need redundancy across replicas and zones, autoscaling that respects slow cold starts, monitoring of latency, queue and GPU memory, a safe rollout process for new weights and engine versions, supply-chain checks on weights, and on-call ownership.
A self-hosted model is a stateful, expensive, slow-to-start service, and it needs the same operational discipline as a database. Redundancy comes first: at least two replicas behind a load balancer, ideally in separate zones, with health checks that test a real short generation rather than only that the process is alive. GPU instances fail, get preempted, or hit driver faults, and a single replica turns each of those into an outage.
Scaling is harder than for stateless web services. A new replica must be scheduled on a GPU node (which may not be available on demand), pull tens of gigabytes of weights, load them, and warm up, which can take several minutes. Keep weights on fast local or regional storage, pre-pull container images, scale on queue time and KV cache utilization rather than CPU, and keep enough warm headroom for the first minutes of a spike. Scaling to zero saves money only where users can tolerate a cold start.
Observability should cover the user view (time to first token, tokens per second, error and timeout rates at p50 and p95), the engine view (queue length, running and waiting sequences, KV cache usage, preemptions, prefix cache hit rate), and the hardware view (GPU memory, utilization, temperature, and error counters). Log prompts and outputs under the same privacy rules as any other user data.
Change management matters because every component moves model behaviour: the weights, the quantization, the engine version, the chat template and the decoding defaults. Pin all of them, version them together, and roll out changes with a canary that compares eval scores and live metrics against the current version before shifting traffic. Keep the previous version deployable for fast rollback.
Security and supply chain round it out. Download weights only from the original publisher, verify checksums, and prefer safe tensor formats over formats that can execute code when loaded, such as Python pickle files. Keep the engine off the public internet behind an authenticated gateway, patch drivers and engine images regularly, and restrict who can change which model is served. Someone must own all of this on call.