Understand Legacy Code Fast With a Structured Prompt
Use a structured understand-legacy-code prompt to map entry points, data flow, and risky areas of an unfamiliar codebase before you change a line.
Getting dropped into a codebase nobody documented is one of the most common jobs in software, and one of the worst served by AI advice. The standard tip is "paste the file into ChatGPT and ask it to explain." That gives you a paragraph that restates the code you already pasted. An understand legacy code prompt does something different: it forces the model to return a structured map of the system, so you know where the entry points are, how data moves, and which areas will hurt you if you touch them.
The difference is the output contract. "Explain this" produces prose. A comprehension prompt produces a fixed shape you can act on, every time, regardless of which model runs it.
This post covers that prompt, the contract that makes it reliable, and the patterns that stop it from hallucinating structure that isn't there.
Why "explain this code" isn't comprehension
Comprehension is building a working mental model of a system: what calls what, where state lives, what breaks if you change a given piece. A paraphrase of one file isn't that. It's a description of syntax you can already read.
The pages that rank for this confirm the gap. DocsBot's legacy codebase documentation prompt does have an output contract with six sections, which is better than most, but it's aimed at producing documentation for monolith decomposition and stays shallow: it extracts "main workflows" and "key attributes" without mapping dependency coupling, data-flow hotspots, or the risky areas that actually bite during a change. Community advice on using AI to understand a codebase lands on "give as much detail as possible and keep the conversation going," which gets you maybe 60% of the understanding but no repeatable structure. There's no copyable prompt whose output is a comprehension map with a locked shape. That's the gap this fills.
What you can do with this prompt
- Get the real entry points of a service, not the
main()you'd guess at. - Trace how a request or a piece of data flows from input to storage to output.
- Find the load-bearing modules: the files that, if you change them, ripple everywhere.
- Surface the risky areas: tangled state, missing tests, silent error handling.
- Produce a "read these five files first" reading order for onboarding.
- Generate the list of open questions to ask whoever still remembers the system.
Anatomy of the prompt
The structure that works feeds the model the code slices and the goal first, then locks the comprehension shape at the end. The output contract is the whole point, so it lives where the model weights it most.
Variables
{{code_context}} → the files/dirs you can paste, plus the repo tree
{{your_goal}} → why you're reading this (fix a bug, add a feature, migrate)
{{known_facts}} → anything you already know (the framework, the entry URL)
Prompt
Role: senior engineer onboarding to an unfamiliar system.
Task: build a structured comprehension map, not a summary.
Context: the variables above.
Output contract (locked)
1. Entry points: where execution starts, with file:line
2. Data flow: input → processing → storage → output, as a path
3. Key modules: the load-bearing files and why each is load-bearing
4. Dependencies: external services, libraries, and shared contracts
5. Risk map: areas that are fragile, untested, or surprising
6. Read-first order: the 5 files to read, ranked, with the reason
7. Open questions: what the code can't tell you, to ask a human
1. Gather inputs
You rarely have the whole repo in one paste. Give the model the directory tree plus the files you suspect matter (the router, the main service class, the config). The tree alone lets it reason about structure even for files it can't see in full.
2. Fill the variables
Put the code and tree in {{code_context}}. State why you're reading the system in {{your_goal}}, because comprehension is goal-shaped. Reading to fix a race condition needs a different risk map than reading to add a feature. Drop anything you already know into {{known_facts}} so the model doesn't waste output re-deriving it.
3. Run the prompt and check the entry points first
The entry-points section is your hallucination tripwire. If the model lists a file:line that doesn't exist, the rest of the map is suspect. Open one or two of the claimed entry points before you trust the data-flow section.
4. Post-process the risk map
The risk map is a hypothesis, not a verdict. Treat each flagged area as "go look here," not "this is broken." The value is direction. It tells you where to spend your limited attention in a system you don't yet trust.
5. Iterate as you learn
Comprehension is a loop. Read the five files it ranked, then feed what you learned back into {{known_facts}} and re-run on a deeper slice. The second pass, anchored on facts you've verified, is far sharper than the first.
Prompt-craft patterns that keep it honest
Pattern 1: demand file:line citations for every claim
Every entry point, data-flow step, and risk you name MUST cite a file path
and line range from the provided context. If you can't cite it, mark it
"INFERRED — not in provided code" instead of stating it as fact.
This is the single most important pattern. Without it, models invent plausible-sounding architecture. With it, you get a map you can check, and the honest "inferred" tags tell you exactly where to look next.
Pattern 2: separate "what the code does" from "what it should do"
Do NOT suggest improvements, refactors, or bug fixes. This is a comprehension
pass. Describe the system as it is. Improvements come later, in a separate run.
Models love to jump to "you should refactor this." On a comprehension pass that's noise. You can't fix what you don't understand, and the suggestions crowd out the map. Hold the line.
Pattern 3: rank the read-first list, don't just list it
The read-first section is an ordered list of exactly 5 files, ranked by how
much understanding each unlocks. For each, one sentence: what reading it teaches.
A flat unranked list is invalid output.
An unordered "here are the important files" is barely better than the directory tree. The ranking is the judgment you're paying for. Here's a model-behavior note worth keeping: Claude honors a "this is invalid if unranked" instruction reliably, while GPT-4o sometimes returns a tidy alphabetical list unless you restate the ranking requirement on the final line.
The opinion that matters most here: a comprehension prompt should refuse to give advice. Every "explain this code" tool on the market rushes to suggest fixes because that demos well. But understanding and improving are different jobs, and mixing them gives you a worse version of both. Map first. Change later.
Variables you'll set
| Variable | Required | What it is |
|---|---|---|
{{code_context}} | Yes | The directory tree plus the key files you can paste |
{{your_goal}} | Yes | Why you're reading: bug fix, feature, migration, audit |
{{known_facts}} | No | Framework, entry URL, anything you've already confirmed |
Getting started
- Grab the repo tree and the three or four files you suspect are central.
- Write your goal in one line; the risk map shapes itself around it.
- Run the prompt and verify the cited entry points exist before trusting the rest.
- Read the five ranked files in order.
- Feed verified facts back into
{{known_facts}}and re-run on a deeper slice. - Turn the open-questions list into actual questions for whoever knows the system.
- For onboarding that recurs every time someone joins, use the Legacy Code Comprehension Harness Agent Pack instead of rebuilding the contract.
Once you understand the system, the next jobs have their own prompts: a Codebase Onboarding Harness Agent Pack turns the map into a structured onboarding doc, and a repo health scorecard prompt grades the same codebase on test coverage and structure.
The Legacy Code Comprehension Harness Agent Pack does this end-to-end: a {{code_context}} variable feeds a comprehension prompt whose locked output contract always returns cited entry points, a data-flow path, a ranked read-first list, and a risk map, with a companion prompt that converts the map into onboarding notes. It's part of The Complete AI Prompts Bundle, a one-time lifetime license to the whole catalog plus every pack added later, useful if you inherit unfamiliar code more than once.
A map you can re-run beats a wiki page that rots. The codebase changes; the prompt doesn't. For the work that comes after comprehension, the technical-debt triage prompt and the change blast-radius prompt both build directly on the structure this one gives you.
See the legacy comprehension pack →Common questions
What is an understand-legacy-code prompt?
Why not just ask ChatGPT to explain the code?
Does Claude or ChatGPT understand legacy code better?
Get the prompt packs this guide is built on
Ready-to-paste prompts with documented variables and usage guides for ChatGPT, Claude, and Gemini. One-time payment, own it forever.
More prompt guides

Run the Same Refactor Across Many Repos With One Prompt
Changing one function signature inside a single repo is a five-minute job. Changing that same signature across nine repos that all call it, when three of them are owned by other teams and deploy on th…

Keep Docs in Sync With Code: A Prompt That Catches Drift
The README says the function takes three arguments. Someone added a fourth six months ago. Nobody updated the docs, because nothing forced them to. That's documentation drift, and it's the default sta…

Service Dependency Audit Prompt: Map Coupling Risk
Every architecture diagram lies a little. It shows the dependencies someone remembered to draw, not the ones that grew in over three years of shipping. A service dependency audit prompt reads the actu…