Skip to main content
Ai promptsPerformanceCoding agentsClaude

A Performance Regression Prompt to Find the Bad Change

Use a performance regression prompt to turn benchmarks, profiles, and a diff into ranked suspect changes plus a verification plan you can act on.

PPromptsCart Team·October 2, 2026·Updated October 2, 2026·8 min read

A dashboard turns red. The p95 latency jumped 40% sometime in the last twenty commits, and now you're staring at a diff with no idea which line did it. The usual answer is git bisect, which is great at finding which commit and useless at explaining why. A performance regression prompt fills that gap: feed it the before-and-after profiles plus the diff, and it returns a ranked list of suspect changes, each with a verification step to confirm or kill it.

The point isn't to replace your profiler. It's to read the profiler output and the diff together, which is the tedious correlation work that eats an afternoon. The model does the cross-referencing; you verify the top suspect.

This post covers that prompt, the contract that keeps it from guessing, and how it slots in alongside the tools you already use.

Why profiler docs don't close the loop

A profiler tells you where time goes. A diff tells you what changed. The regression lives in the overlap, and nothing on the first page of search results connects the two for you.

The ranking content is all manual tooling. Octocore's performance regressions guide is a solid 2,000-word DevOps walkthrough covering baselines, CI tests, git bisect, and profiler comparison with k6, Pyroscope, and Datadog, but it contains no AI prompt at all. The pandas-maintainer write-up on maintaining performance leans on asv continuous and snakeviz to make a slowdown "jump out," which works when the regression is large and obvious. Research on LLM-based diagnosis from diffs notes that a model given only the commit message and diff can reason about changed lines, but it's an academic regression-test-generation paper, not a copyable diagnosis prompt. So there's profiler tooling, there's bisect, and there's research, but no prompt that turns a profile delta plus a diff into ranked, verifiable suspects. That's the gap.

What you can do with this prompt

  • Cross-reference a before/after profile against the diff to find changed code on the hot path.
  • Get a ranked suspect list instead of re-reading the whole diff line by line.
  • Catch the non-obvious causes: an added N+1 query, a lost cache, a dependency bump that changed an algorithm.
  • Produce a per-suspect verification plan so you confirm before you "fix."
  • Distinguish a real regression from a benchmark that just got noisier.
  • Generate the one-line summary for the incident channel once you've confirmed the cause.

Anatomy of the prompt

The structure feeds the evidence first (profiles, diff) and locks the ranking shape last. Models weight recent tokens, so the output contract sits at the end where it survives a long profile paste.

Variables
  {{profile_before}}  → baseline benchmark/profile output
  {{profile_after}}   → the regressed run, same workload
  {{diff}}            → the change between the two versions
  {{workload}}        → what the benchmark actually exercises

Prompt
  Role: performance engineer triaging a regression.
  Task: rank the changes in the diff by likelihood of causing the slowdown.
  Context: the variables above.

Output contract (locked)
  1. Delta summary: what got slower, by how much, on which call path
  2. Ranked suspects: each change, why it's suspect, confidence (high/med/low)
  3. Hot-path mapping: which suspect touches the function that grew in the profile
  4. Verification plan: per suspect, the exact check to confirm or rule it out
  5. Non-regression note: if the delta looks like noise, say so and why

1. Gather inputs

Run the same workload on both versions. The before and after have to exercise the identical path, or you're comparing apples to a different orchard. Capture both profiles and the diff between the two commits or tags.

2. Fill the variables

Paste the baseline into {{profile_before}} and the slow run into {{profile_after}}. Put the diff in {{diff}}. Describe what the benchmark does in {{workload}}, because a model can't map a profile to a hot path without knowing what was being measured.

3. Run the prompt and read the hot-path mapping first

Skip straight to the hot-path mapping. The strongest signal is a suspect that both appears in the diff and touches the function whose time grew in the profile. That intersection is where you start, regardless of how the model ranked confidence.

4. Verify the top suspect before you fix anything

The verification plan exists for a reason. Revert just that one change, re-run the workload, and confirm the delta closes. Plenty of "obvious" regressions turn out to be a second change masking the real one. Confirm, then fix.

5. Iterate down the list

If the top suspect doesn't close the gap, move to the second. Feed the result back: "reverting X recovered 10% of 40%, what else." The model re-ranks with the new evidence, and you converge fast.

Prompt-craft patterns that keep it grounded

Pattern 1: rank only what's in the diff

Every suspect MUST be a change present in the provided diff. Do NOT speculate
about code that isn't shown. If the diff doesn't explain the delta, say
"diff does not account for the regression" and list what profile evidence
points elsewhere.

This stops the model from inventing causes outside the change set. The honest "the diff doesn't explain it" answer is genuinely useful. It tells you to look at infrastructure, data volume, or a dependency you didn't capture.

Pattern 2: confidence has to be earned by profile evidence

A suspect gets "high" confidence only if it touches a function whose time
measurably grew between {{profile_before}} and {{profile_after}}. Changes with
no profile evidence are "low" no matter how suspicious they look in the diff.

Without this, a scary-looking refactor with no actual profile impact gets ranked above the quiet one-line change that doubled a loop. Tie confidence to measured time, not to how alarming the code reads.

Pattern 3: separate regression from noise

First decide whether the delta exceeds normal run-to-run variance for this
workload. If the after-profile is within noise of the before, output the
non-regression note and stop. Don't manufacture suspects for a non-event.

Chasing benchmark noise burns more time than any real regression. Make the model rule out noise before it ranks anything. Claude follows a "stop and report" branch like this cleanly; GPT-4o sometimes plows ahead and ranks suspects anyway, so restating the stop condition near the end helps.

The opinionated take: don't ask the model to write the fix in the same call. Diagnosis and repair are different jobs, and a model that's already committed to "suspect #1 is the cause" writes a fix that assumes it's right. Confirm the cause with the verification plan first. Then, in a fresh call, fix it. The prompt's job is to hand you a short, ranked, checkable list, not a patch you'll trust blindly.

Variables you'll set

VariableRequiredWhat it is
{{profile_before}}YesBaseline benchmark or profiler output
{{profile_after}}YesThe regressed run on the same workload
{{diff}}YesThe change between the two versions
{{workload}}YesWhat the benchmark actually exercises

Getting started

  1. Run the identical workload on the good and bad versions and capture both profiles.
  2. Grab the diff between the two commits or tags.
  3. Describe the workload in one line so the model can map the profile to a path.
  4. Run the prompt and read the hot-path mapping before the confidence ranks.
  5. Revert the top suspect, re-run, and confirm the delta closes.
  6. If it doesn't, feed the result back and walk down the list.
  7. For a repeatable triage flow, the Perf Regression Harness Agent Pack wires the contract into a workflow instead of a one-off paste.
Browse the coding prompt packs →

A performance regression is a triage job, and it has cousins. The Regression Triage Harness Agent Pack handles functional regressions with the same ranked-suspect approach, and a CI failure diagnosis prompt applies the read-the-evidence-then-rank pattern to broken pipelines.

Skip the setup

The Perf Regression Harness Agent Pack does this end-to-end: a {{profile_after}} variable plus the diff feed a triage prompt whose locked output contract always returns a hot-path mapping, profile-backed confidence ranks, and a per-suspect verification plan, with a companion prompt that drafts the incident summary once the cause is confirmed. It's part of The Complete AI Prompts Bundle, a one-time lifetime license to the whole catalog plus every pack added later, worth it if you chase regressions more than once a quarter.

Get the Perf Regression Harness Agent Pack →

Bisect finds the commit; this prompt explains the line. Run them together: narrow the range with bisect, then hand the diff and the profiles to the prompt for the why. For the verification habit that makes any of this trustworthy, the verify AI coding agent output guide and the flaky test detection prompt both reinforce the same confirm-before-you-trust discipline.

See the perf regression pack →
FAQ

Common questions

What is a performance regression prompt?
It's a reusable prompt that takes your before-and-after benchmarks or profiles plus the diff between the two versions, and returns a ranked list of which changes most likely caused the slowdown, each with a verification step to confirm or rule it out.
Can AI find which commit caused a slowdown?
It can rank suspects, not bisect for you. Given the profile delta and the diff, a model reasons about which changed code sits on the hot path and ranks likelihood. You still verify the top suspect, but you start from a short list instead of the whole diff.
How is this better than git bisect?
Bisect finds the commit; it doesn't explain why. A prompt that reads the profile and the diff tells you which line in that commit hits the hot path and how to confirm it. The two combine well: bisect narrows the range, the prompt explains the cause.
Stop reading. Start shipping.

Get the prompt packs this guide is built on

Ready-to-paste prompts with documented variables and usage guides for ChatGPT, Claude, and Gemini. One-time payment, own it forever.