The GenAI Field Guide

Data and synthetic data

Create useful examples without mistaking generated data for ground truth.

Every GenAI system rests on data that someone chose: the examples in a few-shot prompt, the cases in an eval set, the pairs used to fine-tune a model, the labels that train a classifier sitting in front of an LLM. Models have made that data cheap to produce. A single prompt can write a thousand support tickets, label ten thousand emails overnight, or invent invoices in layouts nobody has seen yet. What models have not made cheap is knowing whether the data is any good. Generated text is fluent by construction, so it looks plausible whether or not it is correct, varied, or representative of what real users send.

The mental model for this chapter is real data is the ground truth, everything else is a hypothesis about it. Synthetic examples, model labels and paraphrased test cases are all useful, but each one carries the generating model's blind spots and biases, and each one can leak into places it should not be: a training set that overlaps the test set, or a public benchmark that ended up in a model's pretraining corpus. Most of the craft is measurement: checking generated data against held-out real data, measuring label agreement, finding duplicates, and watching for the day production traffic stops looking like the data you built on.

The Basic questions cover what synthetic data is, why teams generate it, whether to trust it, and how labeling works with people and with models. The Advanced questions cover the ways data quietly corrupts results (leakage, benchmark contamination, near-duplicates), how to validate generated examples and build instruction datasets, when human annotation is non-negotiable, how to measure annotator agreement, what happens when models train on model output, and how to detect drift once the system is live.

What is synthetic data?

Synthetic data is data produced by a program or a model rather than collected from real events: generated support tickets, invented invoices, paraphrased questions or simulated conversations that resemble the real thing closely enough to test or train a system.

Synthetic data is not new. Engineers have long used fake names, random transactions and simulated sensor readings for testing. What changed with LLMs is that generated data can now be unstructured and realistic: a model can write a convincing angry email about a late refund, a contract clause with an unusual indemnity, or a multi-turn chat where the customer changes their mind halfway through.

There are three broad ways to make it. Rule-based generation fills templates from lists or distributions (a name, an amount, a date) and is predictable but repetitive. Model-based generation prompts an LLM to write examples, usually seeded with attributes such as persona, intent, length and difficulty so the outputs vary. Transformation starts from real data and changes it: paraphrasing questions, translating them, swapping entities, or injecting typos and noise. Transformation keeps the real distribution's shape while multiplying volume, which is why it is often the safest option.

Generated data inherits the generator's habits. LLM-written tickets tend to be politely worded, grammatically clean, and average in length, while real tickets are terse, misspelled, angry, or contain a pasted stack trace. That gap is the central risk: a system tuned or tested on clean synthetic data can look excellent and then stumble on the first week of real traffic.

So treat synthetic data as a tool with a specific job (filling coverage gaps, protecting privacy, bootstrapping before real data exists) rather than as a replacement for real examples. The useful question is never "is it synthetic?" but "does it improve or correctly predict performance on real data?"

Why generate synthetic data?

Teams generate synthetic data to cover rare or dangerous cases, to test before real traffic exists, to avoid exposing personal data, and to multiply scarce training examples. It fills gaps in real data; it should not stand in for it.

The most common reason is coverage. Real data follows a long tail: 80% of support tickets might be password resets and delivery questions, while the cases that break a system (a refund dispute in two languages, an invoice with three currencies, a prompt injection hidden in an attachment) appear a handful of times a month. Generating those deliberately lets you test them now instead of waiting to collect them.

The second is cold start. A new product has no logs. Synthetic conversations let you build an eval set, exercise the pipeline end to end and catch format bugs before any user arrives. Once real traffic flows, it should gradually replace the synthetic set.

The third is privacy and access. Real medical notes, bank statements or HR cases may be impossible to share with a vendor or an offshore test team. Synthetic records with the same structure let people build and test without touching the real data. This is not automatic privacy protection, though: a model asked to paraphrase real records can copy names and numbers straight through, so generated sets still need checking.

The fourth is training volume. Fine-tuning a small model, or distilling a large model's behaviour into a smaller one, often uses a strong model to write thousands of input and output pairs. This works well for narrow tasks such as extraction or classification. Check the provider's terms before using one model's outputs to train another, because some restrict it.

The trade-off in every case is that generated data reflects what the generator thinks the world looks like. It fills the gaps you can describe; it cannot reveal the failure modes you have not imagined. Only real traffic does that.

Can synthetic data be trusted?

Only conditionally. Synthetic data is trustworthy for a purpose once you have checked it for correctness, realism, diversity, bias and leakage, and shown that it predicts or improves performance on held-out real data. Unchecked, it mostly measures the generator.

Trust depends on what the data is for. Synthetic inputs for a load test need only the right shape and length. Synthetic cases in an eval set need correct expected answers, otherwise you grade the system against mistakes. Synthetic pairs in a training set need correct outputs and a realistic input distribution, otherwise the model learns the generator's errors and style. The bar rises with how much the data influences decisions.

There are four common failure modes. Wrong labels: the generator writes a question and an answer, and the answer is subtly wrong, so roughly one error in twenty can be enough to teach or reward a mistake. Unrealistic distribution: generated inputs are cleaner, more polite and more uniform than real ones. Mode collapse: hundreds of examples share the same opening phrase, structure or entity names because the model gravitates to its most likely outputs. Inherited bias: if the generator associates certain names with certain jobs or regions, the synthetic set encodes that.

The strongest evidence is downstream. If adding synthetic training data improves accuracy on a held-out set of real examples, it is helping. If a synthetic eval ranks two system versions in the same order as a real eval does, it is a usable proxy. If neither check has been run, the data has not earned trust yet, however good it looks.

Human review remains the cheapest reality check. A domain expert reading 50 random generated examples usually spots the tells in minutes: invoices with impossible tax totals, medical notes no clinician would write, customers who explain their problem too clearly.

What is data labeling?

Data labeling is attaching the correct answer to each example: a category, extracted fields, a reference response or a quality score. Labels define what correct means for training and evaluation, so their consistency limits how good any model can appear.

A label is whatever a grader or a training process needs to know about an example. For a ticket router it is a category (billing, technical, account). For extraction it is the field values (invoice number, total, due date). For a chatbot it might be a reference answer, a list of facts the answer must contain, or a preference between two responses. Labels can also be spans (which words are the drug name) or scores (this summary is a 3 out of 5 for faithfulness).

The work starts before anyone labels: you need a labeling guide that defines each category, gives positive and negative examples, and says what to do with ambiguous cases. Without one, two careful people will disagree. Is "I was charged twice and now I can't log in" billing or account? The guide must decide, for instance by a primary-intent rule, otherwise the dataset contains both answers and the model is penalised whichever it picks.

Label quality sets a ceiling. If annotators agree only 85% of the time, a model cannot reliably show more than roughly that accuracy against their labels, because the remaining cases have no stable truth. That is why teams measure inter-annotator agreement and revise the guide when it is low, rather than labeling more data.

Labeling is iterative. The first 100 examples nearly always expose categories that overlap, are missing or need splitting. Labeling a small batch, reviewing disagreements, updating the guide and relabeling is far cheaper than discovering the problem after 10,000 labels.

Can an LLM label data?

Yes. With a clear rubric and a validation step, an LLM can label most routine examples close to human quality at a fraction of the time, but you must measure it against human labels and route uncertain, novel or high-stakes items to people.

LLM labeling works because many labeling tasks are reading comprehension against a written guide, which is exactly what instruction-following models do well. Give the model the same labeling guide humans use, ask for a structured output (label plus a short reason), and run it over the dataset. For clear-cut categories this often matches human agreement levels.

The catch is that you cannot know it worked without measuring. Label a few hundred examples with people first, run the model on the same set, and compute agreement per category. Models tend to be strong on frequent, well-defined classes and weaker on rare classes, sarcasm, domain jargon and anything requiring context the guide does not state. Per-class numbers reveal that; overall accuracy hides it.

Then decide routing. A practical pattern is prelabel and review: the model labels everything, people review a random sample for quality control plus every item the model flags as uncertain. Confidence can come from asking for a self-rated confidence, from running the model several times and checking whether the label is stable, or from two different models disagreeing. Self-reported confidence is often poorly calibrated, so calibrate it against the human set before trusting a threshold.

Watch for circularity. If an LLM labels the eval set and a similar LLM is being evaluated, they share biases and the eval rewards agreement with the labeler rather than correctness. Keep a human-labelled core for anything used as ground truth.

What is training-data leakage?

Leakage is when information that should be unavailable gets into training or into the inputs: test examples in the training set, future data used to predict the past, or private data that the model can later reproduce. It inflates scores and can expose secrets.

The term covers two related problems. Evaluation leakage makes results look better than they are: the model is tested on things it effectively saw during training. Privacy leakage is when sensitive data enters training and the model can later reproduce it, for example a customer's account number appearing in a generated reply.

Evaluation leakage has several routes. Direct overlap: the same example sits in both splits. Near-duplicate overlap: two copies of one contract with different file names are split randomly, so the test copy is practically memorised. Entity overlap: tickets from the same customer, or pages from the same document, land in both splits, so the model learns that customer's phrasing. Temporal leakage: training uses data from after the period you test on, so the model knows outcomes it could not have known. Prompt leakage: few-shot examples in the production prompt are copied from the eval set, which quietly turns those cases into open-book questions.

Privacy leakage comes from training on raw logs, support transcripts or documents that contain personal or confidential data. Fine-tuned models can memorise rare strings, especially ones that repeat, and reproduce them when prompted with similar context. Redacting before training and testing for memorisation afterwards are both necessary, since redaction misses things.

The fix for evaluation leakage is mostly about how you split: split by customer, document or time rather than by row, and deduplicate across splits before you split. The fix for privacy leakage is to control what enters training at all.

What is benchmark contamination?

Benchmark contamination is when a model's training data included a benchmark's test questions or answers, so its score measures memory rather than ability. Public benchmarks published online are especially exposed, which is why headline scores need your own uncontaminated eval beside them.

Large models are pretrained on enormous web crawls, and public benchmarks live on the web: in their own repositories, in papers, in blog posts that quote them, in forums discussing answers. Unless a training pipeline actively filters them out, test items get absorbed. Even with filtering, paraphrased or translated copies slip through. The result is that a model can score well because it has seen the answer, not because it can solve the problem.

Contamination comes in degrees. Verbatim contamination means the exact question and answer were in training. Partial contamination means the questions without answers, or answers discussed in prose. Indirect contamination means training on data generated by another model that had memorised the benchmark, or on solutions to closely related problems. All of them inflate scores, and none are fully visible from outside because most providers do not publish their full training data.

Detection methods are imperfect but useful. N-gram overlap checks look for long shared sequences (13-grams were used in some early large-model reports) between the benchmark and the training corpus, which requires access to the corpus. Canary strings are unique random markers placed in benchmark files; if a model can complete the marker, the files were in its training data. Perturbation tests rephrase questions, reorder options or change numbers: a model that drops sharply on lightly modified versions was likely relying on memory. Time-based benchmarks use problems published after a model's training cutoff, which is the most robust approach for things like coding contests.

For practitioners the lesson is about how to read numbers. Public leaderboard scores are a rough first filter, not evidence that a model will do your task. Your own eval, built from your data and never published, is uncontaminated by construction. Keep it private, refresh part of it periodically, and avoid pasting its cases into public tools that may train on inputs.

The same mechanism applies inside a team: if engineers repeatedly tune prompts while looking at the eval set, the prompt becomes contaminated by the eval. A held-out slice nobody looks at protects against that.

How do you validate synthetic examples?

Validate in layers: hard constraint checks, deduplication and diversity measures, model or human review of correctness and realism, and finally a downstream test showing the synthetic data improves or predicts results on held-out real examples.

No single check is enough because generated data fails in different ways at different levels. A pipeline that runs cheap checks first and expensive ones last lets you discard most bad examples automatically and spend human attention only where it matters.

Layer one: constraints. These are programmatic and cheap. Does the JSON parse? Does the invoice total equal the sum of line items plus tax? Is the date valid? Is the answer one of the allowed labels? Is the length within the range seen in real data? For generated question and answer pairs, can the answer be found in the source document it claims to come from? Constraint checks routinely reject a noticeable share of generated examples, and every rejection is one fewer wrong label.

Layer two: duplicates and diversity. Remove exact and near-duplicates, then measure spread. Useful signals are the distribution of attributes you asked for (did the 10 intents actually come out evenly?), the number of distinct opening phrases, and embedding-based measures such as average pairwise distance or how well generated examples cover clusters of real ones. If 40% of generated tickets start with "I hope this message finds you well", the set is narrower than its size suggests.

Layer three: correctness and realism. A model judge with a rubric can check whether the expected answer is right and whether the input is plausible, but it shares blind spots with the generator, especially if they are the same model. Sample for human review: domain experts reading 50 to 100 random items, plus every item the judge flags. A simple realism test is to mix real and synthetic examples and see whether a reviewer can tell them apart; if they easily can, the system under test will notice the difference too.

Layer four: downstream utility. This is the decisive test. Train or tune with and without the synthetic data and score both on held-out real data. For eval sets, check that the synthetic eval ranks system versions the same way a real eval does. Data that passes every earlier layer but does not move real results is not worth keeping.

How do you avoid duplicate data?

Normalise and hash records to remove exact copies, use MinHash or embedding similarity to find near-duplicates, then group related items so copies of one source never span the training and evaluation splits. Do this before splitting, and recheck whenever data is added.

Duplicates do two kinds of damage. Within training data, they overweight whatever is repeated: research such as the 2021 paper Deduplicating Training Data Makes Language Models Better found that removing duplicates reduced verbatim memorisation and made training more efficient. Across splits, they leak test answers into training and inflate scores. The second problem is the one that misleads teams most, because the score looks great.

Exact duplicates are easy once text is normalised: lowercase, collapse whitespace, strip boilerplate such as email signatures and timestamps, then hash. Without normalisation, two copies of the same ticket differing only in a trailing space are treated as different.

Near-duplicates need similarity. MinHash with locality-sensitive hashing estimates the Jaccard similarity of word or character shingles (overlapping chunks of five words, say) and scales to millions of documents; a threshold around 0.8 is a common starting point, tuned by inspecting pairs near the boundary. Embedding similarity catches paraphrases that share meaning but few words, which matters for synthetic data where a model rewrote the same example several ways; cosine similarity above roughly 0.95 is a typical first threshold, but it depends on the embedding model, so calibrate it by reading pairs.

Grouping handles the cases similarity misses. Versions of one contract, pages of one manual and turns from one conversation may not look alike but share an origin. Assign each record a group id (source document, customer, conversation) and split by group so a group sits entirely in one split.

Order matters: deduplicate and group first, then split, and keep the dedup index so new data is checked against every existing split, not only against itself. A test set that was clean on day one becomes contaminated when next month's batch is appended carelessly.

How do you build instruction datasets?

Collect representative real tasks, write or curate high-quality responses that follow one consistent style guide, deliberately include edge cases and correct refusals, review for errors, then deduplicate, split and version the set. Quality and coverage matter far more than raw size.

An instruction dataset is a set of prompt and response pairs (or multi-turn conversations) used for supervised fine-tuning: teaching a model to respond to a type of request in a particular way. Work such as the 2023 paper LIMA: Less Is More for Alignment showed that around a thousand carefully chosen, high-quality examples can shape behaviour substantially, which matches practitioner experience: a few hundred to a few thousand excellent examples usually beat tens of thousands of mediocre ones for a narrow task.

Inputs should mirror production. Start from real requests (logs, tickets, user research), redacted, and cluster them to see the task mix. Then fill gaps deliberately: ambiguous requests, missing information, out-of-scope asks, adversarial phrasing, long inputs, other languages. Methods like Self-Instruct generate new instructions from a small seed set; they help with breadth but drift toward generic tasks, so steer them with your own categories.

Responses carry the behaviour you are teaching, so every error is learned. Write a style guide first: length, tone, structure, when to ask a clarifying question, when to refuse and how, how to cite sources. Responses can be written by experts, drafted by a strong model and edited by experts, or selected from several model drafts. The model will imitate exactly what it sees, including hedging, verbosity or wrong formats, so consistency across examples matters as much as correctness. Include abstentions: examples where the right answer is "I don't have enough information" or a handoff, or the tuned model learns to always answer.

Then treat the set as an artefact. Deduplicate, split by source or group, hold out an eval set built the same way, record provenance for each example (real, synthetic, edited, and by whom), and version it. When the fine-tuned model misbehaves you need to trace the behaviour back to the examples that taught it.

When is human annotation required?

Human annotation is required when errors are costly, when the task needs domain expertise or judgement a model lacks, when you need trusted ground truth to measure everything else, and when model labels have not been shown to agree with experts.

The question is less "humans or models?" than "where must the truth come from a person?". In most mature pipelines, people produce a smaller amount of high-trust data and models extend it. Four situations put people in the loop for certain.

High-stakes labels. Medical coding, legal clause classification, fraud decisions and safety ratings have asymmetric costs: a wrong label can produce a model that harms someone or creates liability. Regulated domains may also require documented human review. Here, experts label or verify every item used for training or evaluation, even if a model drafts first.

Expertise and judgement. Models are weaker on specialist jargon, local regulation, rare categories, and tasks whose rules live in people's heads rather than in a guide. They also struggle with subjective quality judgements where the standard is a company's own taste. If the labeling guide cannot fully specify the answer, a model will fill the gap with plausible defaults.

Ground truth for measurement. Every automated label source, including LLM labelers and LLM judges, has to be calibrated against something. That something is a human-labelled set. Without it you cannot say whether the model labeler is 95% or 75% accurate, and you cannot detect it getting worse after a model update.

Novelty and drift. When new categories, products or document formats appear, there are no examples in any guide. People identify the new pattern and define it; models can label it afterwards.

Human annotation is not automatically correct. Annotators get tired, interpret guides differently and drift over long sessions. Measure their agreement too, use expert adjudication for disagreements, and keep a small set of known-answer items mixed into their queue to catch quality drops.

What is data drift?

Data drift is a change in the inputs a live system receives compared with the data it was built and evaluated on: new topics, formats, languages or user behaviour. Quality can fall silently, so you detect it by monitoring input distributions and re-sampling production into evals.

A system is tuned and tested against a snapshot. The world then moves: a new product launches and tickets about it appear, a supplier changes its invoice template, a marketing campaign brings users from a new country, users learn to phrase requests differently, or a regulation changes what counts as a correct answer. None of this raises an error. The model still returns fluent output, and quality falls without anyone noticing.

It helps to separate types. Input drift (also called covariate shift) means the distribution of inputs changes, such as more Spanish tickets or longer documents. Concept drift means the correct answer for the same input changes, such as a refund policy update that makes last quarter's reference answers wrong. Upstream drift means something before the model changes: a new OCR engine, a retrieval index rebuild, or a provider model update, so the same request produces different context or behaviour.

Detection uses signals at several levels. Simple statistics are cheap and catch a lot: input length, language mix, share of each predicted category, rate of empty retrievals, refusal rate. Embedding drift compares where new inputs fall in embedding space relative to the evaluation set, for example the share of inputs far from any eval cluster. Outcome signals such as thumbs-down rate, escalation rate or manual correction rate are the closest to real quality but lag. Classic tabular metrics such as the population stability index can be computed over any categorical or binned feature.

Detection is only useful if it feeds back. The standard loop is to sample recent production traffic every week or month, label a slice, add it to the eval set, and compare scores against the old snapshot. That turns "something might have drifted" into a measured quality change and gives you fresh cases to fix it with.

What happens when models are trained on model-generated data?

Training repeatedly on model output, with real data replaced rather than added to, can cause model collapse: rare patterns disappear, outputs grow more uniform and errors compound. Keeping real data in the mix and filtering synthetic data hard largely prevents it.

When a model generates text it favours its most probable outputs. Rare phrasings, unusual facts and minority styles are under-sampled. If the next model is trained on that output, it learns a slightly narrower distribution; if its own output then trains another model, the distribution narrows again. Research published in 2024, AI models collapse when trained on recursively generated data, showed this progressively erases the tails of the original distribution and, in later generations, produces degraded and repetitive output. This is called model collapse.

The effect depends heavily on setup. Follow-up work found that when synthetic data is accumulated alongside the original real data rather than replacing it, collapse is largely avoided. Much successful practice, such as distilling a large model into a small one for a narrow task or generating training data for specific skills, uses one generation of synthetic data, filtered strictly, mixed with real data and checked against real held-out results. That is very different from recursive self-training on unfiltered output.

Smaller versions of the same problem appear in everyday work. A fine-tune on synthetic responses inherits the generator's verbosity and favourite phrases. A classifier trained on LLM labels inherits the labeler's systematic errors and then looks accurate when evaluated against more LLM labels. A team that generates eval cases with the same model it is testing builds an eval that the model is unusually good at.

There is also a wider data question. As more web content is model-generated, future crawls contain more synthetic text, and filtering it is hard. For teams building datasets, the practical response is provenance: record which examples are real, which are generated and by what, so you can control the mix and remove batches if they prove harmful.

How do you measure label quality and annotator agreement?

Have two or more annotators label the same sample and compute chance-corrected agreement (Cohen's kappa for two, Krippendorff's alpha for more), check accuracy on gold items with known answers, and read the disagreements. Low agreement usually means the guide, not the people, needs fixing.

Raw percent agreement is misleading when classes are imbalanced. If 90% of tickets are billing and both annotators label everything billing, they agree 90% of the time while learning nothing. Cohen's kappa corrects for this by comparing observed agreement with the agreement expected by chance given each annotator's label frequencies. A kappa of 1 is perfect agreement, 0 is chance level. Common rough guidance treats above about 0.8 as strong and below about 0.6 as a sign the task or guide needs work, though acceptable levels depend on how subjective the task is. Krippendorff's alpha generalises to more than two annotators, missing labels and ordinal scales such as 1 to 5 ratings.

Agreement measures consistency, not correctness. Two annotators trained on the same flawed guide can agree perfectly and both be wrong. Gold items, examples whose answers were settled by an expert panel, are mixed into the queue to measure accuracy and to catch an annotator whose quality drifts during a long session.

Per-class agreement is more useful than one number. Disagreement usually concentrates in a few category pairs (billing versus account, minor versus moderate severity). A confusion matrix between annotators shows exactly which boundary is unclear, which tells you which guide rule to rewrite or which categories to merge.

The same tools evaluate model labelers and LLM judges: treat the model as one more annotator and compute its agreement with the human consensus. If the model agrees with experts about as well as experts agree with each other, it is performing at the task's natural ceiling. Expecting more than that is unrealistic, because the remaining disagreement is ambiguity in the task itself.