Structured Output Prompt: Reliable JSON From Any LLM
Write a structured output prompt that returns clean JSON every run, on Claude, ChatGPT, or Gemini. Copy the prompt-only contract and the model behavior table.
Paste a customer review into an LLM and ask for "the sentiment and key topics as JSON," and it works. Do it a thousand times in a pipeline and it breaks: a stray "Here's the JSON you requested," a markdown fence you didn't ask for, a field that's a string one run and an array the next. A structured output prompt exists to stop that drift. It pairs the task with a locked output contract so the same prompt returns the same JSON shape on every call.
Search the topic and most of what ranks is one of two things. Either it's an API tutorial that tells you to flip on OpenAI's strict mode and call it done, like the walkthrough at genaiunplugged.substack.com, or it's a generic "add examples and use delimiters" tip list. Both are useful. Neither hands you a vendor-neutral prompt you can paste into Claude, ChatGPT, or Gemini today, with a behavior table showing where each model needs a nudge.
That's the gap. Not everyone wants to wire up function-calling and a JSON-schema library. Plenty of people just need the paste-and-run version that holds up.
What a structured output prompt actually is
A structured output prompt is a prompt with two halves: a task, and an output contract that names the exact keys, their types, and the rule for missing values. The contract is the part people skip, and it's the part that decides whether the JSON is reliable.
Here's what teams reach for it to do:
- Extract fields from a document (invoice, resume, contract clause) into a fixed record
- Classify text into a set of labels plus a confidence score
- Turn a messy chat transcript into a structured summary a database can store
- Normalize scraped or user-submitted data into one canonical shape
- Score something against a rubric and return the scores as numbers, not prose
- Feed the output straight into another prompt or a downstream API without a parsing step
The common thread: something else consumes the output. A human reading prose doesn't care about a stray sentence. A JSON.parse() two lines later does.
Anatomy of the prompt
Three blocks. Variables for what changes, the task, and the contract last.
Variables
{{input_text}} - the raw text to process (paste the document, review, transcript)
{{schema}} - optional: an inline description of the fields you want
Task
Extract the following fields from the input text.
If a value is not present, use null. Do not infer or guess.
Output contract (state on the LAST line)
Respond with ONLY a JSON object, no prose, no markdown fence:
{
"field_a": string,
"field_b": number | null,
"labels": string[],
"confidence": number // 0.0 to 1.0
}
Two details carry the reliability. First, the null rule: without it, models invent plausible-looking values for fields that aren't in the input, which is worse than a blank because it looks correct. Second, the contract sits on the last line. That's not a style choice. It's where the model's attention is heaviest by the time it starts generating.
Step-by-step usage
1. Write the schema before the prompt
Decide the exact keys and types first, on paper. sentiment as a string from a fixed set ("positive" | "negative" | "neutral") is reliable. sentiment as "however the model wants to phrase it" is not. Pin every field to a type and, where it's a category, to an allowed set of values.
2. Add the missing-value rule
State it once, plainly: "If a value is not present in the input, use null. Do not infer." This single line kills the most common failure mode, which is confident fabrication. A model that can't find the invoice date will happily produce one that looks right. null is the honest answer, and you have to ask for it.
3. Put the contract on the last line
Move the JSON shape to the end of the prompt, after the pasted {{input_text}}. If your input is long, the contract at the top gets diluted by the time generation starts. On the last line, it doesn't. This is the single highest-leverage fix when a prompt that used to work starts drifting on longer inputs.
4. Parse, and fail loud
Wrap the output in a real parser and validate against the schema. When parsing fails, don't silently retry forever: log the raw output so you can see whether the model added a fence, added prose, or produced a genuinely malformed object. Each of those has a different fix, and you can't tell which without the raw text.
5. Feed failures back into the contract
When the same field breaks repeatedly, the fix lives in the contract, not in a post-processing hack. If the model keeps returning "N/A" instead of null, add "use null, never the string 'N/A'" to the rule. The prompt becomes more precise each time it fails, and the parser gets simpler.
How the three models handle it
The behavior differs enough that a prompt tuned only on one model can surprise you on another. Here's what holds up in practice.
| Model | Format adherence | What it needs | Common quirk |
|---|---|---|---|
| Claude | High with a clear schema | A "respond with only JSON" line; no API JSON toggle needed | Will add a friendly preamble unless you forbid it explicitly |
| ChatGPT (GPT-4o) | High, highest with strict mode | Contract restated on the last line; API strict mode for hard guarantees | Wraps output in a ```json fence unless told not to |
| Gemini | Good, improving | An explicit schema and an example output | More likely to reorder keys; pin order in the contract if it matters |
The one instruction that helps all three: forbid anything outside the JSON. "Output nothing but the JSON object. No explanation, no code fence, no trailing text." Claude honors that reliably. GPT-4o needs it near the end. Gemini benefits from one example of the exact shape you want.
Worth saying plainly: OpenAI's own docs note that deeply nested schemas degrade reliability. Flatten where you can. A three-level-deep object with optional arrays inside optional objects is a prompt that'll break on edge cases no matter how you word it.
Skip the six-paragraph "you are a meticulous data extraction assistant" persona. It doesn't help format compliance. A two-line role plus a hard output contract beats a long backstory every time, and it leaves more of the context window for the actual input. The contract does the work, not the character sketch.
Prompt-craft patterns that lock the format
Show one example, in the exact target shape
An example output beats a paragraph describing the output. But the example has to be in the precise shape you want, including how you handle nulls and empty arrays. A model that sees "tags": [] in your example learns to use an empty array; one that sees "tags": "none" learns to improvise.
Forbid the code fence (or demand it) explicitly
The single most common "invalid JSON" bug is a stray json fence the parser chokes on. Pick one: either "wrap the JSON in a json fence" and strip it in code, or "output raw JSON with no fence." Say which. Leaving it unstated is how you get both on alternate runs.
Restate the format after long inputs
For prompts that paste a long document into {{input_text}}, add a one-line reminder right after the input: "Now return the JSON object per the contract above." It's redundant on short inputs and load-bearing on long ones. Cheap insurance.
Variables you'll set
| Variable | Required | What it is |
|---|---|---|
{{input_text}} | Yes | The raw text to process, pasted verbatim: a document, review, transcript, or record |
{{schema}} | No | An inline field list if you'd rather pass the shape as a variable than hardcode it in the contract |
{{allowed_labels}} | No | The fixed set of category values, when a field is a classification rather than free text |
The contract itself isn't a variable. It's the fixed part of the prompt, which is the point: the shape stays constant while the input changes.
Getting started
- Write your exact JSON schema first: keys, types, and any allowed value sets.
- Add the missing-value rule ("use
null, don't infer"). - Put the full contract on the last line of the prompt, after
{{input_text}}. - Add one example output in the precise target shape.
- Test on your three hardest inputs, not your three easiest.
- Wrap the output in a parser that validates against the schema and logs raw output on failure.
- When a field drifts, tighten the contract instead of patching the parser.
For a version of this that goes past a single extraction, where the same context and contract get reused across a whole workflow, the Context Engineering Harness pack is built for it.
Browse the prompt catalog →The Context Engineering Harness Prompt Pack does this end-to-end: an {{input_text}} variable feeds prompts with a locked JSON output contract and an explicit missing-value rule, so extraction returns the same shape whether you run it on Claude, ChatGPT, or Gemini. It's part of The Complete AI Prompts Bundle, a one-time lifetime license to the whole catalog plus every pack added later, worth it if you run more than one of these data jobs.
A locked output contract is one half of reliable prompting. The other half is what you put around it. For assembling the full context an extraction prompt needs, context engineering for LLMs covers the assembly recipe. And when the output feeds a system you depend on, detecting LLM hallucinations shows how to check the model didn't quietly invent a field. Pair either with the Golden Path Scaffolding System Prompt when the JSON is scaffolding a new project rather than filling a database.
Common questions
What is a structured output prompt?
How do I get an LLM to always return valid JSON?
Why does my JSON prompt work sometimes and break other times?
Get the prompt packs this guide is built on
Ready-to-paste prompts with documented variables and usage guides for ChatGPT, Claude, and Gemini. One-time payment, own it forever.
More prompt guides

Cross-Stack Error Correlation Prompt: One Root Cause
A user reports a checkout failure. The frontend shows a failed . The backend logs a 500 with a database timeout. Infrastructure shows a pod that restarted two minutes earlier. Three signals, three das…

AI Coding Agent ROI Prompt: Measure Real Impact, Not Lines
Ask most teams whether their AI coding agent is paying off and you'll get a number from a vendor dashboard: lines accepted, suggestions shown, an "AI adoption rate." None of those tell you whether any…

Golden Path Scaffolding Prompt: Generate a Paved-Road Project
Most teams don't lack project templates. They lack a project skeleton that matches their conventions: the directory layout the platform team blessed, the lint rules CI actually enforces, the test fold…