Change a model's behavior with examples when prompting is not enough.
Fine-tuning is how you change a model's default behaviour with examples instead of instructions. It continues training an existing model on data that shows the inputs you send and the outputs you want, so a format, a house style, a classification scheme or a narrow skill becomes reliable with a short prompt. It is also how teams distil a large model's behaviour into a smaller, cheaper, faster one, or adapt an open-weight model they must run on their own hardware.
The mental model to keep is that fine-tuning shapes behaviour, not knowledge. Weights are a poor database: facts learned in training are recalled unreliably, cannot be cited or permission-checked, and go stale. Retrieval and tools supply what the model must know; fine-tuning shapes what it must do. Fine-tuning also creates a new artefact you own, with a dataset to maintain, evals that must beat a strong prompted baseline, a serving path and a retraining plan.
The questions build in that order. The basic group covers what fine-tuning is, when it is and is not the right tool, how it relates to RAG, what supervised fine-tuning and LoRA actually do, and where synthetic data fits. The advanced group covers QLoRA, how much data you need, catastrophic forgetting, why facts are hard to add, how to evaluate a tuned model, how to prepare data, which hyperparameters matter, the choice between hosted and self-managed tuning, and how to deploy and maintain the result.
Fine-tuning continues training an already-trained model on your own examples, nudging its weights so a specific behaviour (a format, a style, a task) becomes the default without needing to be spelled out in every prompt.
A pre-trained model has learned general language and world patterns from a huge corpus. Fine-tuning takes that model and runs more gradient descent on a much smaller, targeted dataset: usually hundreds to tens of thousands of examples of the inputs you will send and the outputs you want back. The model's weights shift a little so that, for inputs like yours, the outputs you showed it become more probable.
The key mental model is that fine-tuning changes behaviour far more reliably than it adds knowledge. It is very good at teaching a model to always emit a particular JSON shape, to follow a house tone, to classify into your 40 categories, or to do a narrow task with a shorter prompt. It is a poor and expensive way to load in facts that change, because the model stores them diffusely, recalls them unreliably, and you must retrain every time they change.
There are several flavours. Supervised fine-tuning (SFT) trains on input and target-output pairs. Preference tuning (DPO and similar) trains on pairs of better and worse responses. Reinforcement fine-tuning optimises against a grader or verifiable reward. Each can update all the weights (full fine-tuning) or only a small set of added weights (parameter-efficient methods such as LoRA). Hosted model providers expose some of these through a fine-tuning API; with open-weight models you run them yourself.
The trade-off is ownership. A fine-tuned model is a new artefact: it needs a training dataset you maintain, an eval suite that proves it beats the base model plus a good prompt, a serving path, and a plan for retraining when the base model is retired or your requirements change. That overhead is worth it when the same behaviour is needed millions of times; it rarely is for a prototype.
Fine-tune when a stable, repeated task needs consistent behaviour that good prompting cannot reach, or when you need a smaller, cheaper or faster model to match a larger one, and you have quality examples to show it.
The strongest signal is a plateau: you have a measured eval set, you have iterated the prompt, added examples and retrieval, and quality has stopped improving short of your target. If failures cluster in consistent ways (wrong format on edge cases, inconsistent label boundaries, tone drift), examples baked into the weights often fix what instructions cannot, because the model learns the pattern from thousands of demonstrations instead of a paragraph of description.
The second signal is economics. A long prompt with many few-shot examples costs input tokens on every call and adds latency. Fine-tuning moves those examples into the weights, so a short prompt does the same job. It also lets you distil a large model's behaviour into a smaller one: label data with the large model, review it, and train the small model on it. At high volume, a small tuned model can be much cheaper and faster than a large general one for the same narrow task.
The third signal is control. Some teams need a model they can run on their own hardware, at the edge, or in an air-gapped environment, and an open-weight model fine-tuned on their task is the only way to reach acceptable quality at that size.
All of this assumes the task is stable. Fine-tuning freezes a behaviour; if the labels, schema or policy change every month, you will be retraining constantly. And it assumes you can get or create a few hundred to a few thousand examples that are correct, consistent and representative. Without those, fine-tuning just teaches the model your data's mistakes.
Avoid fine-tuning when the real problem is missing or changing information, poor retrieval, a fixable prompt, too little good data, or requirements that are still moving; each of these has a cheaper fix that fine-tuning would only mask.
The most common mistake is fine-tuning to fix a knowledge problem. If the model gives the wrong refund policy, the fix is to retrieve the current policy and put it in context, not to train last quarter's policy into the weights. Facts in weights are hard to update, hard to cite and recalled unreliably, and a tuned model can become more confident about outdated answers.
The second is fine-tuning to fix a retrieval problem. If the RAG system pulls the wrong chunks, a model trained to answer well from good chunks will still answer badly from bad ones. Measure retrieval recall separately; when the right passage is not in the context, no amount of tuning the generator helps.
Third, many failures are prompt problems that look like model limits: contradictory instructions, missing examples of edge cases, no output schema, or a task that should be split into two steps. These take hours to fix; a fine-tune takes weeks and hides the root cause inside weights you cannot read.
Fourth, avoid it when requirements are unstable or data is thin. Early in a product the schema, categories and tone change weekly, and every change invalidates the training set. If you only have 50 examples, or examples you have not checked, you will teach the model noise. Finally, avoid fine-tuning when you need behaviour that can be switched per customer or per request: a prompt or a retrieved policy can change instantly, while a weight change cannot.
There are also governance reasons. Training data may contain personal or confidential information that becomes hard to remove once learned, and a tuned model may need its own review before use in regulated settings.
RAG supplies information at request time from a source you can update; fine-tuning changes how the model behaves. Use RAG for facts that change or must be cited, fine-tuning for consistent format, style or task skill, and often both together.
Retrieval-augmented generation (RAG) searches a knowledge store for passages relevant to the request and puts them in the prompt. The model's weights never change; its answer is grounded in whatever you retrieved. Updating knowledge means updating the index, which can happen in minutes, and every answer can point to its source. Access control is enforceable because you decide per user which documents are retrieved.
Fine-tuning changes the weights so a behaviour becomes the default. It does not give the model a reliable lookup table; facts learned this way are blended into billions of parameters, recalled inconsistently, and cannot be cited or permission-checked. What it does well is make the model consistently follow a pattern: answer in your structure, use your terminology, classify by your rules, or handle a narrow task with a short prompt.
The two are complementary rather than rivals. A common production pattern is a model fine-tuned to use retrieved context well: to quote the passage, to say when the context does not contain the answer, and to answer in the house format, combined with a RAG pipeline that supplies current content. The fine-tune makes the model a better reader; RAG gives it something current to read.
Cost and latency differ too. RAG adds retrieval latency and input tokens on every call but has no training cost. Fine-tuning has an up-front data and training cost, may shorten prompts, and adds the burden of hosting or managing a custom model. When unsure, start with RAG and a good prompt: it is faster to build, easier to debug and easier to change.
Supervised fine-tuning (SFT) trains a model on example conversations where the desired response is known, minimising the error on predicting those response tokens so the model learns to produce similar outputs for similar inputs.
SFT is the simplest and most common kind of fine-tuning. Each training example is an input (a system prompt plus a user message, or a whole multi-turn conversation) and a target response written or approved by someone who knows what good looks like. Training uses the same next-token prediction loss as pre-training: the model sees the input, predicts the response one token at a time, and its weights are adjusted to make the correct tokens more likely.
Two details matter more than people expect. First, loss masking: you usually compute the loss only on the assistant's response tokens, not on the prompt, so the model learns to answer rather than to reproduce your system prompt. Second, the chat template: instruction-tuned models expect special tokens marking system, user and assistant turns. Training data must be formatted with the same template the model will see at inference, or the tuned model behaves unpredictably. Hosted fine-tuning APIs usually handle both if you supply data in their documented message format.
SFT teaches by imitation, so the model learns everything in the targets, including mistakes, inconsistencies and length habits. If half your examples end with a sign-off and half do not, the model will do both at random. If reviewers wrote long answers for hard cases, the model learns that long means careful. Data quality and consistency therefore matter more than volume.
SFT is also the first stage of most assistant training pipelines: providers apply SFT to a base model before preference tuning or reinforcement learning. In application work it is usually the whole job, and preference tuning is added only when "better" is easier to judge than to write.
LoRA (low-rank adaptation) freezes the original model weights and trains small pairs of low-rank matrices added to selected layers, cutting trainable parameters to typically under one percent and making fine-tuning far cheaper in memory and storage.
Full fine-tuning updates every weight, which needs memory for the weights, their gradients and the optimiser state. With the common Adam optimiser in mixed precision that is roughly 16 bytes per parameter before activations, so a 7-billion-parameter model needs over 100 GB of GPU memory. LoRA, introduced in the 2021 paper LoRA: Low-Rank Adaptation of Large Language Models, avoids most of that.
The idea is that the change fine-tuning makes to a weight matrix tends to be low-rank: it can be captured by far fewer numbers than the matrix holds. So instead of updating a large matrix W directly, LoRA keeps W frozen and learns two thin matrices, A and B, whose product B times A has the same shape as W. With a rank r of 8 to 64, the adapter for a 4,096 by 4,096 matrix has around 65,000 to 520,000 parameters instead of 16.8 million. B starts at zero, so training begins exactly at the original model. The output is W x plus a scaled B A x, where the scale is a hyperparameter called alpha divided by r.
Because only the adapters train, gradients and optimiser state are tiny, and the saved artefact is often tens to a few hundred megabytes rather than a full model copy. You can keep one base model and many adapters (one per customer or task), and either merge an adapter into the weights for zero added latency or load adapters dynamically at serving time.
The trade-off is capacity. For format, style and narrow-task adaptation, LoRA usually comes close to full fine-tuning. For large shifts, such as a new language or a heavy domain, full fine-tuning can learn more. A 2024 study titled LoRA Learns Less and Forgets Less found exactly that pattern, which is also LoRA's benefit: it tends to damage the base model's general abilities less.
Synthetic training data is examples generated by software or a model rather than collected from real use, used to fill gaps, cover rare cases or bootstrap a dataset; it helps only when it is filtered, checked and grounded in real inputs.
Real examples are often scarce, expensive to label or too sensitive to use. A capable model can generate candidate inputs (customer messages, documents, queries), candidate outputs (answers, labels, extractions), or both. Rule-based generators and templates can also produce structured variations, for example invoices with different layouts, currencies and date formats.
The most reliable pattern is real inputs, generated outputs, verified: take real or realistic inputs, have a strong model produce outputs following your guide, then check those outputs with rules, a second model, or human review before training. This is how most distillation works: a large model labels data, and a smaller model is trained to imitate it. The opposite pattern, where a model invents both questions and answers from nothing, tends to produce data that is cleaner, more generic and more repetitive than real traffic.
That gap is the core risk. Generated data inherits the generator's style, blind spots and errors. It tends to underrepresent the messy reality of production: typos, mixed languages, incomplete information, angry users. A model trained on it can score well on similar synthetic tests and disappoint on real traffic. Repeatedly training models on model-generated text with little fresh real data has been shown to narrow output diversity over generations, a degradation sometimes called model collapse.
Check the generating model's terms of use as well. Some providers restrict using their outputs to train competing models, and the rules vary by provider and change over time.
QLoRA trains LoRA adapters on top of a base model whose frozen weights are stored in 4-bit precision, cutting memory enough to fine-tune large models on a single GPU, at the cost of slower training and a small quality risk.
LoRA removes most of the gradient and optimiser memory, but the frozen base weights still have to sit in GPU memory. In 16-bit precision that is 2 bytes per parameter: about 14 GB for a 7B model and about 140 GB for a 70B model. QLoRA, from the 2023 paper QLoRA: Efficient Finetuning of Quantized LLMs, stores those frozen weights in 4 bits instead, roughly a quarter of the memory, while keeping the trainable LoRA adapters in 16-bit precision.
The paper introduced three techniques. NF4 (4-bit NormalFloat) is a data type whose 16 levels are spaced to match the roughly normal distribution of neural network weights, so it loses less information than plain 4-bit integers. Double quantisation also quantises the per-block scaling constants, saving a further fraction of a bit per parameter. Paged optimisers move optimiser state between GPU and CPU memory to survive memory spikes from long sequences. During the forward and backward pass, each 4-bit block is dequantised to 16-bit on the fly; gradients flow through the frozen weights into the adapters, which are the only thing updated.
The headline result was fine-tuning a 65B model on a single 48 GB GPU with quality matching 16-bit LoRA on their benchmarks. In practice QLoRA makes 7B to 14B models trainable on a single consumer or workstation GPU and 70B-class models trainable on one data-centre GPU.
The costs are real but modest. Dequantising on every step makes training slower than 16-bit LoRA, often noticeably. Quality is usually close but not guaranteed equal, and depends on the model and task. There is also a deployment subtlety: the adapters were trained against the quantised weights. Merging them into a 16-bit copy of the base model usually works, but you should evaluate the exact artefact you serve, quantised or merged, rather than assume equivalence.
Enough is when adding more examples stops improving your held-out eval. For format and style that is often a few hundred clean examples; for nuanced classification or a hard narrow skill, thousands. Measure it with a learning curve rather than guessing.
There is no universal number because the answer depends on how far the target behaviour is from what the base model already does, and how varied the inputs are. Teaching a capable model to always answer in a fixed structure is a small nudge; a few hundred consistent examples often suffice. Teaching it 60 fine-grained categories with subtle boundaries needs enough examples per category, including the confusable ones, which can mean several thousand. The 2023 LIMA paper made the general point vividly: about 1,000 carefully curated examples produced a strong general assistant, because the base model already had the knowledge and the data only had to shape the style.
Quality beats quantity. A thousand correct, consistent, diverse examples routinely beat ten thousand noisy ones, because SFT imitates everything in the targets. Label noise teaches the model to be inconsistent, and near-duplicates spend training steps re-learning the same case while the model overfits to its phrasing.
Coverage beats volume. What matters is that the training set spans the input variety you will see: every category, the long tail of formats, the hard and ambiguous cases, and the inputs where the right answer is to refuse or ask for clarification. A dataset of 5,000 examples where 4,000 are the easy majority case teaches mostly the easy case.
The way to answer the question for your task is a learning curve. Hold out a fixed, representative eval set. Train on 25%, 50% and 100% of your data with the same settings and plot the score. If the curve is still rising steeply at 100%, more data of the same kind will help. If it has flattened, more of the same will not; you need different data (harder cases, missing categories), cleaner labels, or a different approach. The curve also tells you whether the gain over the prompted baseline is worth the effort at all.
Catastrophic forgetting is when training a model on a new task degrades abilities it had before, such as general reasoning, instruction following or safety behaviour, because the same weights are repurposed for the new data.
Neural networks store many skills in overlapping weights. When you fine-tune, gradient descent moves those weights to fit your examples and has no reason to preserve anything the examples do not exercise. If your dataset is 5,000 terse JSON extractions, the model may get excellent at extraction while getting worse at multi-step reasoning, following unusual instructions, writing fluent prose, or refusing harmful requests. The term comes from 1980s and 1990s research on sequential learning in neural networks, and it applies directly to LLM fine-tuning.
Forgetting is worse when the learning rate is high, when training runs for many epochs, when the dataset is narrow and repetitive, and when all weights are updated. It is particularly visible in safety behaviour: several studies have shown that fine-tuning an aligned model, even on benign data, can weaken its refusals, because the alignment was itself a thin layer of training that new gradients can erode.
Whether forgetting matters depends on how the model is used. A model that only ever performs one extraction task can safely lose general chat ability. A model that also answers free-form user questions, or that sits in a product where users can type anything, cannot.
Mitigations are well established. Use parameter-efficient methods like LoRA, which change less and so tend to forget less. Use a modest learning rate and few epochs, often one to three. Mix in general data (sometimes called replay): add a portion of general instruction-following or safety examples to your training set so those behaviours keep being reinforced. And above all, measure it: keep a regression eval of general capabilities and safety prompts and run it on every candidate model alongside your task eval.
Partly and unreliably. A model can absorb some facts from training, but it needs many varied exposures, recalls them inconsistently, cannot cite them, and may hallucinate more. For facts that matter or change, use retrieval or tools.
Facts do enter models through training; that is how pre-training produces a model that knows capitals and chemistry. The question is whether a small fine-tune is an efficient, reliable way to add your facts, and the evidence says mostly no. Pre-training sees popular facts thousands of times in different phrasings. A fine-tuning set might mention your new product's warranty term twice. Seeing a fact once or twice, in one phrasing, rarely produces reliable recall.
Research supports this. A 2024 study, Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?, found that models learn examples containing genuinely new knowledge much more slowly than examples consistent with what they already know, and that as they eventually fit those new examples, their tendency to hallucinate increases. Other work on the reversal curse showed that a model trained on "A is B" often cannot answer "what is B?", so facts learned in one direction are not stored as a queryable record.
There are practical problems too. You cannot tell which facts the model actually absorbed without testing each one. You cannot update one fact without retraining. You cannot cite a source or enforce permissions on knowledge in weights. And when the fine-tuned fact is wrong or out of date, the model states it with the same confidence as everything else.
What fine-tuning does well adjacent to knowledge is teach domain fluency: vocabulary, abbreviations, the shape of good answers, which details matter. That can make a model better at using retrieved documents in a specialist field. Teams that do want facts in weights, for example for offline or latency-critical use, typically use continued pre-training on large volumes of domain text plus many paraphrased question and answer variants, and still verify recall fact by fact.
Compare the tuned model against the best prompted baseline on a held-out set of real cases, broken down by category, and also run regression checks for general ability, safety, output length, latency and cost. Ship only if it wins where it matters and loses nowhere important.
The baseline is the first decision, and the most often fudged. The fine-tune should be compared not with the base model and a lazy prompt, but with the best prompted system you could build without training: a strong prompt, few-shot examples, maybe a larger model. Many fine-tunes that look impressive against a weak baseline turn out to add little against a good one.
The task eval must be held out properly. Split by the unit that matters in the real world (customer, document template, time period), not by random row, so near-identical cases cannot sit in both train and test. Use real examples, and if possible a time-based split, training on older data and testing on newer, so you measure how the model handles drift. Report per-category results and look at examples, because averages hide regressions in rare but important cases.
Then run regression evals for what the fine-tune might break: general instruction following, safety and refusal behaviour, behaviour on out-of-scope inputs, and the ability to say "I don't know". Track operational metrics too: average output length (fine-tunes often drift longer or shorter), format validity rate, latency and cost per task. A model that is 3 points more accurate but twice as verbose may be a net loss.
Training loss is not an evaluation. Watch validation loss during training to choose a checkpoint and spot overfitting, but make the release decision on task metrics. Where outputs are open-ended, use pairwise comparison (tuned versus baseline, judged by people or a calibrated LLM judge, with position swapped) rather than absolute scores. Finally, confirm in production with shadow traffic or a canary before full rollout, because the eval set is never quite the real distribution.
Collect real, representative examples; write a labelling guide and fix targets to follow it; remove errors, personal data and near-duplicates; format with the model's chat template; and split train and test by a real-world unit so nothing leaks across.
Start from the production distribution. Pull inputs from real traffic or the closest realistic source, and sample deliberately so rare categories and hard cases are present. Then make the targets consistent: write a short labelling guide covering format, tone, length and how to handle ambiguous or out-of-scope inputs, and rewrite or relabel examples that break it. Inconsistency in targets is the single most common reason fine-tunes underwhelm.
Clean aggressively. Remove examples with wrong answers (spot-check samples; if more than a few percent are wrong, the source needs work). Strip or replace personal and confidential data unless you have a clear basis to train on it, because models can reproduce training text verbatim. Remove boilerplate that you do not want learned, such as legal footers or agent signatures. Normalise formats so dates, currencies and whitespace look the way you want the model to produce them.
Deduplicate at two levels. Exact duplicates are easy. Near-duplicates (the same template with different numbers, or a paraphrase) are the dangerous ones: within training they cause overfitting to one phrasing; across splits they leak test cases into training. Use normalised hashing or embedding similarity to find them, then keep each cluster in one split.
Split by the unit that generalises. If the model must handle new customers, split by customer; new document templates, by template; future traffic, by time. Keep the eval set fixed and versioned so results are comparable across runs. Finally, format records with the exact chat template and system prompt you will use in production, include examples where the right answer is a refusal or a clarifying question, and record the dataset version alongside every trained model.
Learning rate, number of epochs and the data itself dominate; for LoRA, rank and which layers are adapted come next. Most other settings have reasonable defaults. Tune against held-out validation loss and task metrics, not training loss.
Learning rate controls how far each update moves the weights, and it is the setting most likely to make or break a run. Too high and the model forgets general abilities, overfits quickly or becomes unstable; too low and it barely changes. LoRA usually needs a higher rate than full fine-tuning because it updates far fewer parameters. Common starting points in open-source recipes are around 1e-4 to 2e-4 for LoRA and around 1e-5 to 2e-5 for full fine-tuning, with a short warmup and a decaying schedule. Hosted fine-tuning APIs often expose a multiplier instead of an absolute rate.
Epochs, the number of passes over the data, are the second lever. With small datasets, one to three epochs is typical. Watch validation loss: when it stops falling and starts rising while training loss keeps dropping, the model is memorising your examples. Save checkpoints and pick the one with the best validation metrics, which is often not the last.
For LoRA, rank sets adapter capacity and target modules set where adapters go. Rank 8 to 32 covers most application tasks; adapting all linear layers usually beats attention-only adapters for similar cost. Alpha scales the update and interacts with the learning rate, so change one at a time.
Batch size (often reached through gradient accumulation) trades stability against steps per epoch; larger batches usually tolerate a slightly higher learning rate. Maximum sequence length must cover your longest real examples or they will be truncated silently, often cutting off the very response you wanted the model to learn. Packing, combining short examples into one sequence, speeds training but needs correct attention masking so examples do not see each other.
Hyperparameters rarely rescue bad data. If a sensible default run underperforms, look at the examples and the eval before launching a large search.
A provider's fine-tuning API is the fastest path and needs no GPU work, but you do not get the weights and are bound to their methods, models and deprecation schedule. Tuning open weights yourself gives full control and portability in exchange for running training and serving.
Hosted fine-tuning means uploading a dataset to a model provider, starting a job, and getting back a private model identifier you call through the same API. The provider handles GPUs, training code and serving. What you can tune is limited to the models and methods they offer, which may include supervised fine-tuning, preference tuning and, at some providers, reinforcement fine-tuning against a grader you define. Hyperparameter control is usually coarse. You never receive the weights, so you cannot run the model elsewhere, and when the provider retires the base model you must retrain on its successor. Data handling terms for training data vary by provider and need the same review as any other data you send.
Self-managed tuning means choosing an open-weight model, running training with open-source tooling on your own or rented GPUs, and serving the result yourself or through a host that accepts custom weights. You control every setting, own the artefact, can run it on-premises or air-gapped, and can keep using it as long as you like. In exchange you take on GPU capacity, training reliability, inference serving, scaling, security patching and the expertise to debug all of it.
The choice is usually decided by three questions. Do you need the base quality of a model that is only available through an API? Do data residency, air-gap or portability requirements rule out a hosted service? And does your team have, or want to build, the skills to run training and inference? Volume matters too: at sustained high volume a small self-served model can be cheaper per call, but only once you count engineering time and idle GPU hours honestly.
Some managed platforms sit in between, running training and serving of open-weight models for you while letting you export the weights. These can be a sensible middle path when portability matters but GPU operations do not appeal.
Version every model with its base, dataset and eval results, roll it out behind a canary with a fallback to the prompted baseline, monitor drift, and plan retraining for when data changes or the base model is retired.
A fine-tuned model is a build artefact and deserves the same discipline as code. Record, for every candidate, the base model identifier and version, the dataset version and its hash, the training configuration, the eval results and who approved release. Without that record you cannot reproduce a model, explain a regression or answer an auditor's question about what data shaped it.
For serving, LoRA adapters offer two paths. You can merge the adapter into the base weights and serve it as an ordinary model with no added latency. Or you can keep adapters separate and serve many from one base model: several open-source serving engines can load multiple LoRA adapters and route each request to the right one, which is how teams run per-customer or per-task adapters without a GPU per model. Separate adapters cost a little per-request overhead but make swapping and rollback cheap. Hosted fine-tunes are served by the provider under a model identifier.
Roll out gradually. Shadow the new model on live traffic, then send a small canary share, compare task metrics, format validity, latency and user signals against the current model, and keep a fallback: if the tuned model errors or fails validation, the request can go to the prompted baseline. Treat the prompted baseline as a permanent fallback rather than something you delete.
Maintenance is the cost people forget. Inputs drift: new products, new document layouts, new user phrasing. Monitor per-category quality on sampled production traffic and harvest failures into the next dataset version. The base model will eventually be superseded or retired, which means retraining on a new base and re-running the full eval, not just the task metric. Set a review cadence, and give the model an owner.