Structured output and constrained decoding
The moment an LLM's output stops being prose and starts being a parameter to a program, the failure mode changes shape. A malformed JSON blob is not a slightly worse answer; it is an exception in a downstream service. This part walks the ladder from politely asking for JSON to a decoder that physically cannot emit anything else — and then draws the line that no decoder can cross: a constrained grammar guarantees the shape of the output and says nothing whatsoever about whether it is true.
The ladder to reliable structure
Ask nicely → JSON mode → constrained decoding
There are three rungs and they are not equivalent. The first is instruction and example: put the schema in the prompt, show one well-formed instance, and hope. It works often enough to be tempting and fails often enough to break a pipeline — and the failures cluster on exactly the hard cases: long outputs, deep nesting, unusual values, content that contains quotes or newlines. The second rung is JSON mode, where the API guarantees the output is parseable JSON but not that it matches any particular schema; you still validate, and you still handle retries. The third rung is constrained (or guided) decoding, where the sampler itself is only allowed to consider tokens that keep the output a valid prefix of the target grammar.
Under the third rung sits a finite-state machine built from the grammar. At each decoding step the engine knows which state it is in, computes the set of vocabulary tokens that would keep the string a valid prefix, masks every other logit to negative infinity, and samples from what is left. The mask below is hand-built for the tiny schema {"name": <string>, "age": <int>} over a deliberately small twelve-token vocabulary. Step through it and watch the allowed set change while the emitted text stays a valid prefix.
Top: the grammar's states, left to right. Middle: the text emitted so far. Bottom: the twelve-token vocabulary — highlighted tokens are allowed in the current state; the rest are masked.
How the mask is built
Grammar to FSM, and why nesting needs a stack
Engines like XGrammar (arXiv:2411.15100), Outlines (arXiv:2307.09702) and LLGuidance all take the same shape. Compile the target format — a JSON Schema, a regex, a context-free grammar such as a SQL dialect — into an automaton over the model's vocabulary. At each step, walk the automaton's current state, look at each vocabulary token, and mark it legal only if consuming it keeps the string on a path to an accepting state. Then apply that mask to the logits before sampling. Practical engines precompute per-state masks and cache the small number of states a real request actually visits, which is why constrained decoding is close to free in latency.
JSON needs more than a finite-state machine, and that is the detail most explanations skip. Nesting is not regular: {"a": {"b": {"c": …}}} requires counting that the closing braces match the opening ones, so a plain FSM cannot recognise it. The engines use a pushdown automaton — an FSM plus a stack — where entering an object pushes and the closing brace pops. Vocabulary tokens that span a structural boundary, such as a token that contains both a value and a closing brace, are handled by advancing the automaton over the token's characters one at a time and masking it only if every character is legal.
Two refinements are worth knowing because they change the felt behaviour. Strict function-calling schemas are the hosted version of the same idea: you declare the tool arguments as JSON Schema, the provider enforces it (OpenAI's strict mode, for instance, additionally requires that every property be required and that optionality be expressed through unions with null). Jump-forward decoding goes further: when the automaton reaches a state where the next characters are forced — the ":{" between two known keys — the engine appends those tokens directly instead of sampling them, saving steps and removing a whole class of low-probability mistakes.
Left: seeded parse-success rate against the number of schema fields. Bottom: twenty individual unconstrained outputs at the current setting — red is a parse failure.
xGrammar's headline claim is that its compilation and mask generation are fast enough to add only microseconds to a decode step: the cost of correctness here is engineering, not throughput.
Syntax is not semantics
Valid JSON that says the wrong thing
This is the sentence to carry away from the part. A constrained decoder guarantees that the output parses. It does not guarantee that a single field is correct. Every failure mode of an unconstrained model survives intact inside a well-formed envelope: wrong numbers, transposed dates, hallucinated fields, the wrong enum value, a unit mismatch between the field name and the value. What changed is only that the failure now arrives as a typed object that your program will happily act on, which is frequently worse than a parse error — a crash is loud, and a plausible wrong value is silent.
Notice also what the decoder can do to your confidence. Constraining the output makes the model's tokens near-deterministic in the structural positions, and a downstream reader sees a tidy, syntactically flawless object. The tidiness is manufactured by the grammar, not earned by the content. The demo below makes the point with a real JSON.parse: the object is provably valid, and it contradicts the source on several fields at once.
Left: the source text. Right: the model's output. Red marks flag fields whose values contradict the source, even though the object parses.
Schema design is API design
Every nullable union is a branch your caller must handle
Because the schema is enforced mechanically, it stops being documentation and becomes an interface. Three habits pay off. Prefer an enum to free text whenever the set of answers is closed: "status": {"enum": ["new", "triaged", "closed"]} makes an off-vocabulary value unreachable, where "status": "string" merely invites the model to invent a synonym. Mark fields required unless you have genuinely thought about the absent case, because downstream code that assumes presence fails at the worst moment; the strict-mode rules encode exactly this instinct. And avoid ambiguous unions: a value that may be a string, a number or an object forces every consumer to branch, and it usually means the model is being asked to make a modelling decision it should not own.
The panel below toggles those three decisions and reports a seeded mock validation-failure rate over 300 trials. The baseline is small; free text, optional fields and ambiguous unions move it by tens of points, which is the difference between an extraction pipeline you can trust and one that pages you at night.
Left: seeded downstream validation-failure rate for the current schema and two references. Right: the schema the toggles imply.
When you cannot constrain
Retry, validate, and fail loudly
Constrained decoding is not always available. The output may be prose that must contain a structured fragment; the provider may not support a grammar; the schema may be too large to compile; or the task may be a free-form answer you only intend to post-process. In those cases the discipline is a validation loop with a hard budget: parse, validate against the schema, and on failure re-ask once with the validator's own error message appended to the conversation. That last detail matters — "field total expected number, got string" is a far better repair prompt than the original instruction repeated.
Two rules make the loop safe. First, never let a retry loop run unbounded; three attempts and a typed failure is almost always better than a request that hangs a worker. Second, validate twice: once for the schema and once for the semantics you actually care about — bounds, cross-field consistency, agreement with the source. The first check is what a constrained decoder gives you for free; the second is the check that catches the failure this part exists to warn about.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| What does JSON mode guarantee? | Parseable JSON and nothing more — not your schema, not your field names, not your types. |
| What does a grammar constraint guarantee? | That the emitted token sequence is a valid string of the grammar. Syntax only. |
| Why is an FSM not enough for JSON? | Nesting needs a stack to match braces; the engine uses a pushdown automaton, not a finite-state machine. |
| What is jump-forward decoding? | Appending forced tokens directly instead of sampling them, skipping steps the grammar has already decided. |
| Does constrained decoding improve accuracy? | Not on the values. It removes malformed-output failures and can otherwise leave content quality unchanged or slightly worse. |
| Enum or free text? | Enum whenever the answer set is closed; free text invites synonyms a downstream switch will not match. |
| Optional or required? | Required unless the absent case is designed for; consumers assume presence and fail at the worst moment otherwise. |
| How do you repair a failed parse? | Re-ask once with the validator's error message appended, bounded to a few attempts, then fail loudly. |
Further reading
- Dong et al., "XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models", 2024 — grammar compilation and near-zero-overhead token masking.
- Willard & Louf, "Efficient Guided Generation for Large Language Models", 2023 — the Outlines approach: format constraints as FSM-guided decoding.
- Beurer-Kellner et al., "Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation", ICML 2024 — why masking must be subword-aligned, and what naive masking costs in task accuracy.
- OpenAI, Structured outputs guide, 2024 — strict JSON Schema for tool arguments, including the all-fields-required rule.
- JSON Schema, Specification — the vocabulary the schemas above are written in; draft revisions are dated and stable enough to pin.
- Microsoft Guidance, LLGuidance — a context-free-grammar engine that computes a token mask per step with Earley parsing over a token trie, reports roughly 50 µs per mask, and now backs OpenAI's JSON Schema enforcement.