Make model output safe and predictable for software to consume.
Most valuable LLM features are not conversations. They are steps inside a program: classify this ticket, extract these invoice fields, choose this route, fill this form. Each step needs the model's output in a shape code can read without guessing. Structured generation is the set of techniques that make that happen, from asking politely for JSON, to provider modes that enforce a JSON Schema, to grammars that restrict output to a safe subset of SQL.
The mental model is a contract with two halves. Enforcement, through constrained decoding, guarantees the shape: the right keys, types and allowed values. Validation, in your own code, checks the meaning: does this order exist, is this amount within policy, does this citation say what the answer claims. Enforcement makes parsing reliable, but it can never make a value true, and in one important case (a required field the source cannot fill) it can push the model to invent one. Good systems use both halves and treat model output like input from an untrusted client.
The Basic questions cover what structured output is, why JSON and JSON Schema are the common contract, how constrained decoding works, why validation still matters and when prose is the better choice. The Advanced questions cover what to do when validation fails, grammars beyond JSON, the gap between valid and correct, optional fields, schema versioning, the line between tool arguments and final answers, how field order affects quality, and how to stream structured output safely.
Structured output is a model response in a defined, machine-readable shape, usually JSON matching a schema, so that code can consume named fields directly instead of interpreting free text.
A language model naturally produces prose. Software, on the other hand, needs values it can route, store and compare: a customer_id, an issue_type drawn from a fixed list, an urgency level. Structured output is the practice of making the model return those values in an agreed shape, most often a JSON object described by a schema, so the next step in the program reads result["urgency"] rather than searching a paragraph for the word "urgent".
There are three broad ways to get it, with very different guarantees. Prompt-only: you describe the format and give an example, and the model usually complies, but nothing stops it from adding a friendly preamble or a trailing comma. JSON mode: the provider guarantees syntactically valid JSON, but not which keys or types it contains. Schema-constrained output (often called structured outputs or strict mode): you pass a JSON Schema and the provider restricts generation so the response must match it. Tool calling uses the same machinery, since tool arguments are structured output aimed at a function.
The reason this matters is that most production LLM features are not chat. They are extraction, classification, routing and form filling inside a larger program, and every one of those needs a reliable contract between the probabilistic part and the deterministic part. Without it, teams write fragile regular expressions over prose, and a small model update that changes phrasing silently breaks the parser.
Structure is a contract about shape, not about truth. A response can match the schema perfectly and still contain a wrong customer id or a misclassified issue. That is why the rest of this chapter pairs every format guarantee with validation, business checks and evals.
JSON gives the application named, typed fields it can parse with a standard library, validate against a schema and store directly, instead of guessing values out of prose that changes wording between runs.
The core benefit is that parsing becomes a solved problem. Every language has a JSON parser, every database can store it, and every API speaks it. When a model returns {"sentiment": "negative", "score": 0.82}, the application can read both values in one line, and a type error fails loudly at the boundary instead of propagating a garbled string into a report.
JSON also gives you names. A prose answer like "I think this is mostly negative, maybe 8 out of 10" forces the code to infer which number is which and what scale it uses. Named fields make the meaning explicit, and they let you add a field later without breaking existing consumers that ignore unknown keys.
It pairs naturally with JSON Schema, the standard vocabulary for describing which keys exist, their types and their allowed values. The same schema can drive constrained generation at the provider, validation in your code, documentation for other teams and test fixtures. One artifact serves as the contract for the whole pipeline.
JSON is not free. Braces, quotes and repeated key names cost output tokens, and output tokens are typically priced higher and generated more slowly than input tokens. For a 20-field object with long key names, the syntax can be a meaningful share of the response. Short but clear key names help, and so does asking only for fields the code actually uses. Some teams use other formats such as YAML, XML-style tags or CSV for specific cases, but JSON has the best support for enforcement across providers.
Finally, JSON is for machines. When the consumer is a person reading an explanation, forcing it into a single explanation string field adds escaping and cost for no benefit. The usual pattern is a JSON envelope with the decision fields plus one free-text field for the human-facing message.
JSON Schema is a standard, machine-readable way to describe what a valid JSON document looks like: which properties exist, their types, which are required, allowed values and ranges. It is the contract between a model and the code that reads its output.
A JSON Schema is itself a JSON document. It says, for example, that an object must have a reason field equal to one of four known values, an amount that is a number no lower than zero, and no other properties. Validators exist in every major language, so the same schema can check output in Python, TypeScript or Go.
The building blocks are few. type sets the kind of value (object, array, string, number, integer, boolean, null). properties lists an object's fields and required says which must be present. enum restricts a value to a fixed list. minimum, maximum, minLength and pattern constrain ranges and formats. additionalProperties: false forbids keys you did not declare. description documents a field, and providers pass descriptions to the model, so a good description improves the output as well as the docs.
In LLM work the schema does three jobs. It is passed to the provider to constrain generation, it is used by your code to validate the result, and it documents the contract for other teams. Many teams never write the schema by hand: they define a typed model (a Pydantic class in Python or a Zod schema in TypeScript) and generate the JSON Schema from it, so the type and the schema cannot drift apart.
One important caveat: providers that enforce schemas during generation usually support only a subset of JSON Schema. Some ignore or reject keywords such as pattern, numeric bounds, string formats or certain combinations of anyOf. Some require every property to be listed as required and additionalProperties to be false in strict mode. The subset varies by provider and changes over time, so check the current documentation and enforce anything unsupported in your own validator.
Constrained decoding restricts which tokens the model may choose at each generation step, masking any token that would make the output violate a schema or grammar, so the result is guaranteed to match the required structure.
A model generates text one token at a time. At each step it produces a score for every token in its vocabulary, and the sampler picks one. Constrained decoding inserts a filter between those two steps: before sampling, it sets the probability of every token that would break the format to zero. If the schema says the next value must be a number, tokens like " or abc are simply not available. The model still chooses among the allowed tokens using its own preferences, so the content is the model's while the shape is guaranteed.
To know which tokens are allowed, the system converts the schema or grammar into a state machine and tracks where in that machine the output currently is. The hard part is that tokens do not line up with grammar symbols: one token may contain ": plus the start of a value. Efficient implementations precompute, for each state, which vocabulary tokens are legal, so the mask costs little per step. Some providers compile a new schema on first use, which can add noticeable latency to that first request before the compiled form is cached.
This is what sits behind providers' strict structured-output and strict tool-calling modes, and behind open-source inference servers that accept a JSON Schema, regular expression or grammar. It removes a whole class of failures: missing braces, wrong types, invented enum values and extra keys.
It does not guarantee everything. If the response hits the maximum token limit, you get a valid prefix that is incomplete JSON. If the model would rather refuse, constraints may force it to produce a well-formed but meaningless object, which is why some providers return refusals through a separate channel. And forcing a model down a path it found unlikely can lower quality slightly, especially when the schema leaves no room to think before answering.
Yes. Enforcement guarantees shape, not meaning. A schema-valid object can still hold an id that does not exist, an amount above policy, a date in the wrong year, or a value the user is not authorised to request, so validate in your own code every time.
Think of validation in three layers. Syntax: is it parseable JSON at all? Schema: are the required keys present with the right types and allowed values? Semantics: do the values make sense for this business and this user? Constrained decoding, when it works, covers the first two. Only your code can cover the third, because only your code knows that order ORD-10293 belongs to a different customer or that refunds above 500 need a manager.
Even the first two layers deserve a check in code. Providers enforce a subset of JSON Schema, so bounds, patterns and formats may not be applied during generation. Responses can be truncated by the token limit. A configuration change, a fallback to a different provider, or a model that does not support strict mode can remove enforcement without anyone noticing. A validator at the boundary turns all of those into a clear, logged error instead of corrupt data downstream.
Semantic checks are where the real protection lives. Cross-check extracted values against the source (does the order id actually appear in the email?), against your systems (does the order exist, and does this user own it?) and against policy (is the amount within the refund limit?). Treat the model's output exactly as you would treat input from an untrusted web form, because when the model has read user-supplied or retrieved text, that is effectively what it is.
Validation results are also data. Log every failure with the field, rule and model version. A sudden rise in one rule failing is often the first sign that a prompt change, a model update or a new kind of input has shifted behaviour.
Use prose when the reader is a person who needs explanation, nuance or persuasion, and no code will branch on the content. Use JSON when software must store, route or act on specific values. Many products need both, in one envelope.
The deciding question is: who consumes this, and what do they do with it? If code reads a value and takes a different path depending on it, that value belongs in a typed field. If a person reads the text to understand something, prose is the better medium, because it can carry caveats, ordering and tone that a set of fields flattens.
Forcing prose into JSON has real costs. Long text inside a JSON string needs escaping for quotes and newlines, which wastes tokens and occasionally trips up parsers. Streaming becomes awkward because the user-facing text is buried inside a structure that is incomplete until the end. And some evidence suggests that tight format constraints can reduce answer quality on open-ended reasoning tasks, though how much depends on the model and the schema, and results in published studies are mixed.
Going the other way has costs too. A prose answer that a downstream system must parse is a time bomb: the parser works until the model rephrases. If a support bot must both explain a policy to the customer and tell the ticketing system whether to issue a refund, those are two consumers with two needs.
The common resolution is a hybrid envelope: a small JSON object with the decision fields plus one string field for the human-facing message, or a tool call for the action plus a normal prose reply. Another option for long documents is markdown with a fixed set of headings, which a person reads naturally and code can split reliably on the headings.
Classify the failure first. Retry with the validator's error message when the output is fixable, fall back or escalate to a person when it is not, and never pass invalid data downstream. Cap retries and log every failure.
Validation failures come in kinds, and each kind wants a different response. Parse failures (not JSON, or truncated) usually mean the format was not enforced or the token limit was hit; the fix is to enable strict mode or raise the limit, and a blind retry often fails the same way. Schema failures (missing field, wrong type, value outside the enum) are well suited to a repair retry. Semantic failures (an order id that does not exist, an amount over policy) may mean the input truly lacks the information, in which case retrying invites the model to invent something.
A repair retry sends the original request again with the invalid output and the validator's exact error, for example "delivery_date: '31/02/2026' is not a valid ISO date". Specific error text works much better than "try again", because the model can see what to change. One or two retries usually recover most fixable failures; beyond that, success rates fall and you are paying for the same mistake repeatedly.
When retries are exhausted, the system needs a defined fallback. Options include a stronger model, a simpler schema that captures less, a partial result with the failing fields marked unknown, or a human review queue. Which is right depends on the cost of each error: a misrouted support ticket is cheap, a wrong payment is not. Whatever happens, the invalid object must not reach a database or an action.
Make retries safe. A retry must not repeat side effects, so validate before any tool executes or record is written. Bound total attempts per request so a pathological input cannot loop forever, and include the attempt count and failure reason in traces. A rising retry rate is a quality signal worth an alert, because it often means a prompt or model change made things worse while the retry loop quietly hid it.
Grammar-constrained generation limits output to strings accepted by a formal grammar, such as a context-free grammar for a SQL subset or a domain language, extending constrained decoding beyond JSON to any syntax you can define precisely.
JSON Schema covers JSON. Many applications need other formats: a restricted SQL dialect, a query language for a search engine, a configuration format, a function-call syntax, or a fixed set of command strings. Grammar-constrained generation takes a formal grammar, typically written in a notation like Backus-Naur form (BNF) or as a regular expression, and applies the same token-masking idea: at each step only tokens that keep the output a valid prefix of some string in the grammar are allowed.
The mechanism scales with the grammar's power. A regular expression can be compiled to a finite-state machine, which is cheap to track. A context-free grammar handles nesting, such as balanced parentheses or subqueries, and needs a pushdown automaton (a state machine with a stack). Open-source inference engines and libraries offer both, and some hosted APIs accept a grammar or regex directly; support varies widely by provider, so check before designing around it. Internally, JSON Schema enforcement is usually implemented by converting the schema into exactly this kind of grammar.
The real value is that you can shrink the language to what is safe. Instead of letting a model write arbitrary SQL, the grammar can allow only SELECT statements over named tables and columns, with no semicolons, no comments and no data-modifying keywords. An entire category of dangerous outputs becomes impossible to generate, not just unlikely.
Two limits matter. First, syntax is not semantics: a grammar-valid SELECT can still read a table the user should not see, so authorisation must still happen where the query runs, with a read-only role and row-level security. Second, a grammar the model has rarely seen in training can lower quality, because the model is pushed into token sequences it finds unlikely. Keep grammars close to familiar syntax, and measure task accuracy with and without the constraint.
No. A schema checks the shape and type of values, never whether they are true. A perfectly valid object can carry a hallucinated date, a wrong total or a citation that points at the wrong page, so correctness needs separate checks.
Constrained decoding makes the model's output fit the container. It says nothing about whether the content came from the input or from the model's imagination. Worse, the constraint can increase fabrication in one specific case: when a required field has no value in the source, the model must still emit something of the right type, and a plausible-looking invented value is the easiest way to satisfy the schema.
Well-formed output is also more persuasive. A paragraph with a hedge sounds uncertain, but "invoice_total": 1840.50 looks authoritative in a table and flows straight into a spreadsheet. Teams that would never trust an unchecked prose answer sometimes trust structured output simply because it validated.
Correctness needs checks aimed at meaning. Grounding checks confirm extracted strings actually appear in the source, or that a quoted span exists at the cited location. Cross-field checks catch internal contradictions, such as line items that do not sum to the stated total. System checks look values up: does this customer, product or regulation exist? Evals with labelled examples measure field-level accuracy, which is the number that actually matters, rather than schema pass rate.
Schema design can help. Allow null or an explicit unknown for fields that may be missing, so the model is not forced to invent. Ask for a short evidence or source_span field next to high-stakes values, so a checker can verify each one. For citations, require an identifier your system issued (a chunk id from retrieval) rather than free text, so a fabricated citation fails a lookup instead of looking plausible.
Decide explicitly what absence means, then encode it consistently: usually keep the key required and allow null, so missing information is stated rather than guessed. Make consumers tolerate null, and never let the model fill gaps with invented values.
Optional fields hide an important question: does a missing value mean "not in the source", "not applicable", "the model forgot" or "the extraction failed"? If the contract does not say, every consumer guesses differently, and analytics quietly mix genuine absences with errors.
There are two ways to express optionality in JSON Schema. Leaving a property out of required means the key may be absent. Keeping it required but typed as ["string", "null"] means the key is always present and may be null. For model output, the second is usually better. The model must make an explicit decision for every field, the output shape is stable for consumers, and several providers' strict modes require every property to be listed as required anyway, using a null union to express optionality. Check your provider's rules, since this varies.
When absence has more than one meaning, say so in the schema. A phone field can be null for "not provided", while a separate status or an enum value like "not_applicable" covers other cases. For extraction, some teams pair important fields with a short evidence string so a null can be checked: did the document really not contain a phone number?
The prompt and field descriptions must give permission to be empty: "Use null if the document does not state this. Do not infer." Without that, models tend to fill every slot, because their training rewards helpful, complete answers. On the consumer side, code must handle null deliberately: a display layer shows "not provided", a database column is nullable, and a downstream rule that needs the value routes the record for follow-up instead of crashing. Also watch for empty strings and placeholder values like "N/A" or "unknown" leaking in where null was intended; normalise them at the validator.
Treat an output schema like a public API contract: give it an explicit version, make additive changes backward compatible, and for breaking changes run old and new versions side by side, migrate consumers, then retire the old one.
An output schema has more consumers than people expect: the parser, the database table, analytics jobs, a downstream service, cached results, eval datasets and stored traces. Renaming customer to customer_id looks like a tidy-up, and it breaks every one of them. Schemas for model output deserve the same discipline as a REST API's response format.
Separate changes into additive and breaking. Adding an optional or nullable field is usually safe if consumers ignore unknown keys. Adding a value to an enum is not always safe: a consumer with an exhaustive switch statement will fail on the new value. Renaming, removing a field, changing a type or tightening a constraint are breaking. A useful rule is to be strict about what you generate and lenient about what you accept: the model is held to the current schema, while readers of stored data tolerate older versions.
Make the version explicit. Put schema_version in the stored record (set by your code, not generated by the model), and keep each version's schema, prompt and validator together in source control. A breaking change then follows a parallel run: release v2, write both formats or translate v1 to v2 at read time, move consumers one by one, and remove v1 only when telemetry shows nothing reads it. For stored history, either backfill with a migration function or keep a reader for each live version.
Schema changes also change model behaviour. A new field, a renamed key or a reordered object can shift what the model produces for the other fields, because the schema is part of the prompt. So a schema change should run the regression eval suite exactly like a prompt change, and eval datasets need their expected outputs migrated alongside the schema.
Tool arguments are structured data the model produces to ask your code to do something, such as look up an order. Final structured output is the result the model hands back to the caller. They use similar machinery but carry different risks and checks.
Both are JSON that matches a schema, and providers often enforce both with the same constrained decoding. The difference is direction and intent. A tool call says "please run lookup_order with {"order_id": "ORD-10293"}": it is a request into your system, and your code decides whether to execute it. Final structured output says "here is the OrderSummary you asked for": it is the answer flowing out to whoever called the model.
That difference changes what you check. Tool arguments are the entry point for actions, so they need authorisation (may this user look up this order?), idempotency for writes, rate limits and sometimes human approval. A prompt-injected document that makes the model emit refund(amount=5000) is a tool-argument problem. Final output is the entry point for data, so it needs correctness checks, grounding against what the tools returned, and schema compatibility for downstream consumers.
In an agent loop the two interleave. The model calls tools, reads results, perhaps calls more, and finally produces the answer. A clean design keeps them separate: tools have small, specific schemas for their inputs, and the final answer has its own schema for the caller. Some systems implement the final answer as a special tool (often called something like submit_result) so the loop ends when the model calls it. That is a valid pattern, but treat that tool's arguments with final-output checks, not action checks.
A common mistake is mixing the two: asking for a final JSON object that includes an action field your code then executes. That bypasses the tool layer's authorisation and approval checks, because the action arrives disguised as data.
It can. A model writes fields in order, so a schema that demands the verdict first leaves no room to reason before committing. Putting evidence or reasoning fields before the answer field, and keeping schemas simple, usually recovers most of the loss.
A model generates tokens left to right, and every token conditions on the ones before it. In a JSON object, that means fields are effectively decided in the order they are written. Most constrained-decoding implementations emit properties in the order the schema lists them. If the first field is "approved": true, the model has committed to the decision before writing any of the justification that follows, and the justification becomes a rationalisation of a choice already made.
Ordering fixes much of this. Put fields that gather evidence or reason about the input before the field that states the conclusion: relevant_clauses, then analysis, then decision. The model now does its thinking in tokens the decision can attend to. This is the structured-output version of asking for reasoning before the answer. With reasoning models that think internally before responding, the effect is smaller, because the hidden reasoning happens before any JSON is written, but ordering still costs nothing.
Whether strict formats reduce quality more broadly is contested. A 2024 study titled Let Me Speak Freely? reported lower reasoning accuracy under strict format constraints, while later analyses argued much of the gap came from prompt and parsing choices rather than the constraint itself. The practical reading: effects exist, vary by model and task, and are usually small for extraction and classification and larger for multi-step reasoning. Measure on your own eval set rather than assuming either way.
Other design choices matter too. Deeply nested schemas, dozens of fields and long enums all make the task harder. Field names and descriptions act as instructions, so customer_sentiment with a one-line description beats cs. When a task needs heavy reasoning, a two-step approach is often better: let the model reason in prose (or internally), then produce the structured result in a second, cheap formatting step, or in the same call after the reasoning field.
Parse the JSON incrementally as tokens arrive, render fields to the user as each one completes, and treat the object as untrusted and incomplete until the stream ends and full validation passes. Never act on a partial object.
Streaming improves perceived latency because users see progress after the first token instead of waiting for the full response. Plain text streams naturally. JSON does not: until the closing brace arrives, the text is not valid JSON, and a standard parser will reject every intermediate state.
The solution is a partial JSON parser, which tolerates an unfinished document by closing open strings, arrays and objects in its working copy and returning whatever fields are present so far. Several SDKs and open-source libraries provide this, and some providers stream structured output as incremental deltas you assemble yourself. The UI can then show a title as soon as it is complete, list items as each one closes, and a summary field as it grows.
Schema design affects how well this works. Put fields the user wants first at the top, since generation follows schema order. Prefer arrays of small objects to one huge string, so each item completes and renders independently. Avoid making the first field a long reasoning block if users are waiting to see something, unless you deliberately hide it and show a progress indicator.
Two rules keep streaming safe. First, a partial field value is not a value: a number streamed as 12 may become 1250, and a string may be cut mid-word, so only display fields the parser marks complete, and never write partial data to storage. Second, validation happens once, at the end. Constraints like enum membership, cross-field rules and business checks cannot be judged on a fragment. If final validation fails, the UI must be able to retract or replace what it showed. Any tool calls or side effects wait for the complete, validated object.