Choose a model using your own tasks, not a leaderboard alone.
Picking a model is one of the few decisions in a GenAI system that touches everything at once: answer quality, latency, cost per task, where data may go, and how much engineering you must do around the model to make it reliable. The market moves fast enough that any specific recommendation goes stale within months, so this chapter teaches the method and the mechanisms rather than a ranking. The method is simple to state: define what good looks like on your own work, measure candidates against it under equal conditions, and pick the cheapest option that clears the bar with margin.
The mental model is a frontier of trade-offs. Larger and more heavily reasoning models usually score higher but cost more and respond more slowly; smaller, quantized or distilled models are cheaper and faster but fail more often on hard cases. Most production systems do not sit at one point on that frontier. They route easy requests to cheap models, escalate hard ones, and use serving techniques such as quantization and speculative decoding to move the frontier itself. Public benchmarks give a rough first filter, but they measure someone else's tasks under someone else's settings.
The questions build in that order. The Basic questions cover how to choose, small versus large, routing, cascades, benchmarks, open-weight models and quantization. The Advanced questions explain the mechanisms that shape a model's cost and quality profile (mixture-of-experts, reasoning effort, distillation, speculative decoding), then return to measurement: how to compare models fairly, why leaderboards mislead, and how to switch models safely once you are in production.
Choose an LLM by writing down hard constraints, shortlisting a few candidates that meet them, and scoring each on an eval built from your real tasks for quality, latency and cost per successful task.
Start with hard constraints, because they eliminate candidates before any testing. Where may the data go (a provider API, a specific cloud region, only your own hardware)? What context length do your largest inputs need? Do you need tool calling, structured output, image input or a particular language? Is there a latency ceiling, such as a voice agent that must start speaking within a second? A model that fails a hard constraint is out no matter how well it scores elsewhere.
Next, build a task eval: 100 to 300 real examples of the work, with a way to grade each output. Pull them from logs, tickets or documents rather than inventing them, and include the awkward cases (long inputs, ambiguous requests, the edge cases users complain about). Grade with code where possible (exact match, schema validity, tests passing) and with a calibrated model judge or human review where judgement is needed.
Run each shortlisted model through the same eval with a reasonable prompt for each. Record quality, p95 latency (the latency 95% of requests beat), and cost per task including retries and reasoning tokens. Then compute cost per successful task: a model that is half the price but fails twice as often can cost more once you count retries and human cleanup.
Finally, weigh the soft factors: rate limits and capacity, provider reliability, version pinning and deprecation policy, and how easy it is to switch later. Put the model behind an interface in your code so the choice is reversible. The decision is not permanent, and the eval you built is what lets you revisit it cheaply when new models appear.
Use the smallest model that meets your quality bar on your own eval; large models win on hard, open-ended reasoning, while small models often match them on narrow, well-specified tasks at a fraction of the cost and latency.
Larger models, meaning more parameters and usually more training compute, tend to know more facts, follow complex instructions better and handle multi-step reasoning more reliably. But the gains are uneven. On narrow tasks with clear instructions (classification, extraction into a schema, short rewrites, routing) a small model often scores within a point or two of a large one, because the task does not exercise the extra capacity.
Small models are cheaper per token and faster in two ways: lower time to first token (the delay before output starts) and higher tokens per second. They can run on a single GPU or even on a device, which matters for privacy and offline use. They are also easier to fine-tune, so a small model trained on a few thousand examples of your task can beat a general large model on that task.
Large models earn their cost where errors are expensive or the task is open-ended: planning, ambiguous user requests, long-document synthesis, code changes across many files, and anything where the model must notice what it was not told. They also tend to be more robust to sloppy prompts and unusual inputs, which reduces engineering effort early in a project.
The practical answer is rarely one or the other. Start with a strong model to establish what good looks like and to label data, then test whether a smaller model can meet the bar on each sub-task. Keep the large model for the steps that need it, and route or cascade the rest.
Model routing sends each request to the model best suited to it, using rules, a classifier or a small model to decide, so cheap models handle easy traffic and expensive models only see what needs them.
Real traffic is mixed. In a typical assistant most requests are simple lookups or short rewrites, and a minority need deep reasoning. Sending everything to the strongest model pays the top price for the easy majority; sending everything to the cheapest model fails the hard minority. A router sits in front and picks a model per request, before any generation happens.
Routers come in three forms. Rules use known signals: the product feature that sent the request, input length, the presence of attachments, or the user's plan. A trained classifier (often a small fine-tuned encoder or a logistic regression on embeddings) predicts difficulty or category from the request text. An LLM router asks a small model to label the request. Rules are cheapest and most predictable; classifiers are fast and can be trained on your eval data; LLM routers are flexible but add a call and their own error rate.
The key difference from a cascade is that routing decides up front, from the input alone. That keeps latency low, because each request makes one main call, but it means the router must predict difficulty without seeing an answer. A router that misjudges a hard request as easy produces a bad answer with no second chance, so measure its mistakes directly.
To build one, label a sample of traffic with which model was actually needed (run both models, grade both outputs), then train or tune the router to predict that label. Evaluate the routed system end to end: overall quality, the share of traffic sent to each model, and cost per task. Set the threshold to trade cost against quality explicitly.
A model cascade tries a cheaper model first, checks its answer, and escalates to a stronger model only when the check fails, so you pay for the expensive model only on the requests that need it.
A cascade differs from routing in when it decides. A router predicts difficulty from the input; a cascade looks at the cheap model's actual output and decides whether it is good enough. That makes the decision better informed, at the cost of extra latency on escalated requests, which pay for two calls. The 2023 paper FrugalGPT showed large cost reductions from this pattern on several benchmarks.
Everything depends on the acceptance check. The strongest checks are deterministic: the output parses against the schema, the extracted total matches the sum of line items, the generated code passes its tests, the cited passage actually contains the quoted text. Weaker but still useful signals include token log-probabilities (where the API exposes them), agreement across two or three samples, or a small judge model. A model saying it is confident is not a reliable signal on its own.
The economics depend on the escalation rate. If the cheap model is a tenth of the price and handles 80% of requests, the cascade costs roughly 0.1 + 0.2 = 0.3 of the large model alone, plus the cost of checks. If the cheap model only passes 30% of requests, you pay for both models on 70% of traffic and may spend more than using the large model directly. Measure the pass rate before building.
Cascades can have more than two levels and can end with a human. They work best on tasks with a cheap, trustworthy correctness check, and worst on open-ended writing where judging quality is as hard as producing it.
Benchmarks are fixed, public test sets with a scoring rule, used to compare models on a capability such as knowledge, math, coding or tool use; they are useful for a first filter but rarely predict performance on your own task.
A benchmark is a dataset plus a scoring rule plus a protocol: how the question is presented, how many examples are shown, how answers are extracted and graded. Common families test broad knowledge with multiple-choice questions, grade-school or competition math with checkable final answers, code generation with unit tests, real repository bug fixes judged by a test suite, long-context retrieval, and multi-step tool use in simulated environments.
Benchmarks are valuable because they are shared and repeatable. They let you see broad capability differences, track progress over time, and spot obvious weaknesses, such as a model that does well on knowledge questions but poorly on tool use. For a shortlist, a few relevant benchmark scores narrow dozens of models to three or four worth testing.
Their limits are just as structural. Scores depend on the protocol: the prompt format, number of few-shot examples, sampling settings, reasoning budget and answer parser can move a score by several points, so numbers from different reports are often not comparable. Popular benchmarks saturate, with top models clustered near the ceiling, where differences are within noise. Test questions leak into training data (contamination), inflating scores. And a benchmark measures one narrow skill under clean conditions, while your task mixes skills with messy inputs and specific formats.
Read a benchmark as evidence about a capability, not about your product. A strong math score says something about multi-step reasoning; it says little about whether the model will follow your contract-review rubric or keep to your JSON schema.
An open-weight model is one whose trained weights are published for download under a license, so you can run, inspect and fine-tune it on your own hardware; the license, not the word open, decides what you may do with it.
A model is a set of numbers (the weights) plus code to run them. With a hosted API you send text and get text back; the weights stay with the provider. With an open-weight model you download the weights and run inference yourself, on your own GPUs, a cloud instance, or a laptop for smaller models. That gives you control over where data goes, which version runs, and how long it stays available.
Open-weight is not the same as open source. Most releases publish weights and inference code but not the training data or full training recipe, so you cannot reproduce or fully audit how the model was built. Licenses vary widely: some are permissive (Apache 2.0 or MIT style), others add conditions such as acceptable-use policies, attribution, limits on using outputs to train other models, or special terms above a user-count threshold. Read the license before you build a product on it.
The benefits are control and flexibility: data never leaves your environment, you can fine-tune fully (not just through a provider's tuning API), you can quantize for cheaper hardware, and a version cannot be deprecated out from under you. The costs are operational: you own serving, scaling, monitoring, security patches and GPU capacity. For low or spiky traffic, hosted APIs are often cheaper, because you pay nothing for idle GPUs.
Many providers also host open-weight models behind APIs, so the choice is not binary. You can start on a hosted endpoint for an open-weight model and move to self-hosting the same weights later if volume or data rules justify it.
Quantization stores a model's weights (and sometimes activations or the KV cache) in fewer bits, such as 8 or 4 instead of 16, cutting memory roughly in proportion and often speeding up generation, at some cost in quality.
Model weights are normally stored as 16-bit floating-point numbers, so each parameter takes 2 bytes. A 70-billion-parameter model therefore needs about 140 GB for weights alone, more than any single common GPU holds. Quantization maps those values onto a smaller set of numbers: 8-bit integers or 8-bit floats (1 byte each) halve memory, and 4-bit formats (half a byte) quarter it. The same 70B model at 4 bits needs roughly 35 GB plus overhead.
Speed improves because generating each token is mostly limited by memory bandwidth: the GPU must read every active weight from memory for every token. Fewer bytes to read means more tokens per second, especially at small batch sizes. Some hardware also has fast low-precision math units, which helps prompt processing too.
Quality loss depends on the method and the bit width. 8-bit quantization is usually close to lossless. 4-bit is often acceptable with good methods that calibrate on sample data and handle outlier values carefully (GPTQ and AWQ are well-known examples), but losses show up first on hard reasoning, math, long contexts and less common languages. Below 4 bits, degradation is usually noticeable. Quantizing the KV cache (the stored attention keys and values for the context) is a separate choice that saves memory on long contexts.
Quantization is mainly a self-hosting tool. Hosted providers may quantize internally, but you rarely control or see it. When you quantize, rerun your own eval: average benchmark scores can hide a drop on exactly the cases you care about.
Mixture-of-experts (MoE) is an architecture where each layer holds many parallel sub-networks called experts and a router activates only a few per token, so total parameters are large while compute per token stays small.
In a standard dense transformer, every token passes through every parameter of every layer. In an MoE transformer, the feed-forward block of some or all layers is replaced by a set of experts (for example 64 or 128 feed-forward networks) plus a small router (also called a gate). For each token at each layer, the router scores the experts and sends the token to the top few, often one, two or eight, and combines their outputs weighted by the router scores. Attention layers are usually still shared.
This separates two numbers that are the same in a dense model: total parameters and active parameters. A model might have hundreds of billions of parameters in total but use only a few tens of billions per token. Compute per token, and therefore cost and speed of generation at scale, tracks the active count. Knowledge capacity tracks something closer to the total. That is why MoE models can match much larger dense models in quality at a fraction of the training and inference compute.
The catch is memory. All experts must be loaded, because any token may need any expert, so an MoE model needs memory for its total parameter count. Serving it efficiently usually requires spreading experts across GPUs (expert parallelism) and large batches so every expert gets enough tokens to keep the hardware busy. At low batch sizes or on a single machine the speed advantage shrinks, and memory is the binding constraint.
Training MoE models is harder: routers can collapse onto a few favoured experts, so training adds load-balancing objectives or routing biases, and the token-to-expert assignment introduces communication overhead. Despite the name, experts are not tidy topic specialists. Studies of their routing find specialisation is mostly on token-level and syntactic patterns, not on subjects like law or biology.
For model selection, MoE matters mainly when you self-host. A hosted MoE model is just a model with a price and a speed. Self-hosted, you must budget memory for total parameters while expecting throughput closer to a dense model of the active size, and only if your traffic keeps batches full.
Reasoning effort is a request setting on reasoning models that controls how many hidden thinking tokens the model may spend before answering, trading latency and cost for accuracy on hard problems.
Reasoning models are trained, usually with reinforcement learning on verifiable tasks, to produce a long internal chain of thought before the final answer. The thinking explores approaches, checks intermediate results and backtracks. More thinking generally helps on problems that need search or verification, such as math, debugging, planning and multi-constraint scheduling. This is one form of test-time compute: spending more inference work per question instead of using a bigger model.
Providers expose the dial differently. Some offer named levels (low, medium, high), some take an explicit thinking-token budget, some let the model decide adaptively, and some switch reasoning on or off. The setting is usually a target, not a guarantee: the model may stop early on an easy question or run up to the limit on a hard one. Check whether the reasoning text is returned, summarised or hidden, and whether thinking tokens are billed as output tokens; with most providers they are.
The costs are real. Thinking tokens add latency before the first visible token, often seconds to minutes at high effort, and they can dominate the bill because output tokens are priced higher than input. Gains show diminishing returns: accuracy typically climbs quickly from no reasoning to moderate effort, then flattens. On simple tasks such as extraction or classification, extra effort adds cost with no gain and can occasionally hurt, as the model overthinks or second-guesses a clear answer.
Treat effort as another model-selection axis. A mid-size model at high effort may beat a large model at low effort on your reasoning-heavy task, and cost less. The only way to know is to sweep the setting on your eval and plot accuracy against cost and latency per task.
Use effort per request, not per application. A coding agent can use low effort for reading files and summarising, and high effort for planning a change or diagnosing a failing test.
Distillation trains a smaller student model to reproduce the behaviour of a larger teacher, either from the teacher's generated outputs or from its full probability distributions, so the student gets much of the teacher's quality on a task at lower cost.
The idea dates to the 2015 paper Distilling the Knowledge in a Neural Network. A teacher's output is richer than a hard label: its probabilities across all options encode which wrong answers are nearly right. Training a student to match those soft targets transfers more information per example than training on correct labels alone, so a small student can learn faster and generalise better.
For LLMs there are two common forms. Sequence-level distillation is the practical default: run the teacher on many task inputs, keep good outputs (often filtered by checks or a judge), and fine-tune the student on those input-output pairs with ordinary supervised fine-tuning. It works with any teacher you can call, including hosted APIs. Logit-level distillation trains the student to match the teacher's full next-token probability distribution at every position, usually with a KL-divergence loss. It transfers more signal but needs access to the teacher's logits and a compatible tokenizer, so it is mainly used with open-weight teachers or inside model labs. Many small open models are themselves distilled from larger ones.
Task-specific distillation is where most product teams benefit. If a large model scores 94% on your ticket classifier and a small model scores 81% out of the box, generating 10,000 teacher-labelled tickets and fine-tuning the small model often closes most of the gap. The student is then cheaper, faster and can be self-hosted. It will not match the teacher on tasks outside the distilled distribution, so keep the scope narrow.
Two cautions. The student inherits the teacher's errors and biases, and filtering teacher outputs matters more than volume. And terms of service for some hosted models restrict using outputs to train competing models; read the terms for your use case, and check your own legal position before distilling from a provider's model.
Speculative decoding speeds up generation by letting a cheap draft method propose several tokens that the large target model then checks in a single forward pass, keeping the target model's output distribution while emitting multiple tokens per step.
Generating text is sequential: each token needs a full forward pass of the model, and each pass is limited mostly by reading the weights from memory, not by arithmetic. The GPU has spare compute during decoding. Speculative decoding, described in two 2023 papers from separate research groups, uses that spare compute. A fast draft proposes the next k tokens; the target model scores all k positions in one parallel pass, which costs about the same as generating one token.
The target then accepts draft tokens from left to right as long as they agree with what it would have produced. With greedy decoding, a draft token is accepted if it equals the target's top choice. With sampling, a rejection-sampling rule accepts each token with a probability based on the ratio of target to draft probabilities, and resamples from a corrected distribution at the first rejection. This rule makes the final text statistically identical to sampling from the target alone. Every step also yields at least one token from the target, so the worst case is about normal speed plus the draft overhead.
Speed-up depends on the acceptance rate: how often the draft guesses what the target would say. Predictable text (code, structured output, boilerplate, text that copies from the prompt) gets high acceptance and speed-ups of two to three times are commonly reported. Creative or high-temperature text gets less. Drafts come in several forms: a small model from the same family sharing the tokenizer, extra prediction heads trained on the target (as in Medusa and EAGLE), or n-gram lookup that copies spans from the prompt, which is cheap and effective for editing and RAG answers.
The trade-offs sit in serving. The draft costs memory and compute. At large batch sizes the GPU is already busy, so the spare compute disappears and gains shrink or vanish. Serving engines such as vLLM support several speculative methods; tune k and the method for your traffic. Hosted providers may use it internally, which is one reason output speed differs across endpoints of the same model.
Compare models fairly by running all of them on the same held-out task set with the same grader, equal prompt-tuning effort, matched budgets and repeated runs, then reporting paired differences with confidence intervals alongside latency and cost.
Most unfair comparisons are accidental. The incumbent model's prompt has been tuned for months while the challenger gets the same prompt pasted in. One model runs with reasoning on and the other off. A max-token limit truncates the more verbose model. The eval set was built from failures of model A, so it measures A's weaknesses rather than general quality. Each of these can swing results by more than the real difference between models.
Hold the instrument constant. Same cases, same grader version, same parsing, same tools and retrieval, same timeout and retry policy. Give each model a fair prompt: either a shared, neutral prompt, or an equal fixed budget of prompt iteration per model on a separate development set, never on the test set. Match sampling settings to how you will run in production, and record reasoning effort and token limits as part of the configuration under test.
Account for noise. Model outputs vary between runs, and a 300-case eval has sampling error of a few points. Run each model two or three times or use several samples per case, and compare models paired by case: for each case, did A pass and B fail, or the reverse? A paired bootstrap or McNemar's test gives a confidence interval on the difference. If the interval includes zero, you have not shown one is better.
Report the full trade-off, not one score: quality with interval, p50 and p95 latency, cost per successful task, and failure categories (format errors, refusals, wrong answers). Two models with the same accuracy can fail in very different ways, and one failure type may be cheap to catch while the other reaches users.
Finally, guard against contamination and leakage. Keep a held-out set that no prompt was tuned on, and refresh it periodically from new real traffic.
Leaderboards mislead because their tasks differ from yours, their scores are sensitive to evaluation settings, their test data leaks into training, crowd votes reward style, and the ranking itself creates pressure to optimise for the board.
A leaderboard compresses a model into one number on someone else's distribution of tasks. Even when the measurement is honest, the ordering on that distribution can differ from the ordering on yours. A model that tops a coding board may still be worse at your language, your framework or your repository's conventions.
Contamination is the best-known problem. Public test questions appear on the web and end up in training data, so a model may have seen the answers. Scores rise without capability rising, and performance drops on fresh questions written in the same style. Benchmarks that refresh questions over time, or hold out private test sets, reduce this but do not remove it.
Settings sensitivity is the quieter one. The same model can move several points with different prompt templates, few-shot examples, answer parsers, reasoning budgets or agent scaffolds. Self-reported scores often use the most favourable settings, and different organisations' numbers on the same benchmark are often not comparable.
Preference arenas, where people vote between two anonymous answers, measure what raters like in a quick read, not what is correct. Votes tend to favour longer answers, confident tone and heavy formatting, so arenas often add style controls to correct for this. Few voters check facts or run code. Ranking methods also carry uncertainty: adjacent positions are frequently statistically tied.
Finally, Goodhart's law applies: when a measure becomes a target, it stops being a good measure. Labs can test many private variants and publish the best, tune for the board's format, or choose which benchmarks to report. None of this requires bad faith; it is the natural result of competing on a public number. Leaderboards remain useful for a shortlist and for spotting broad trends. They are not evidence that a model will do your job well.
Switch models like upgrading a critical dependency: pin the current version, run the new candidate through your full eval suite, adapt prompts, shadow or canary it on real traffic, compare cost and latency, and keep a one-step rollback.
Model changes happen whether you plan them or not. Providers release new versions, retire old ones on a published schedule, and sometimes update a model behind an unversioned alias. Each change can shift output format, verbosity, refusal behaviour, tool-calling habits and latency. Prompts tuned for one model often lean on its quirks, so a model that is better on benchmarks can still break your product.
The first defence is pinning: call an explicit, dated model version rather than a floating alias, and record the model id with every logged request. Then a change is a deliberate decision with a diff, not a surprise. Track the provider's deprecation dates so you migrate on your schedule, not in the week before a shutdown.
The second defence is a regression suite you can run on demand: your task eval, format and schema checks, tool-calling traces, safety and refusal cases, and every past production failure turned into a test. Run the candidate through it with the current prompt first, to see what breaks, then adapt the prompt. Expect to change prompts: newer models often follow instructions more literally, need less emphasis, or use tools differently.
The third defence is staged exposure. Shadow testing sends a copy of real traffic to the new model without showing users its output, so you can compare on real inputs. A canary then serves a small share of users and watches quality signals, error rates, latency and cost per task before widening. Keep the old model configured so rollback is a configuration change, not a deploy.
Watch the second-order effects: a more verbose model raises output token cost and latency, a model that calls tools more often changes downstream load, and a stricter refusal policy can raise support tickets. Measure cost per successful task before and after, not only accuracy.