An OKR Drafting Prompt That Catches Vanity Key Results
A reusable OKR drafting prompt that locks one objective to three measurable key results and flags vanity metrics. Includes model behavior and a contract.
An OKR drafting prompt has to fight the model's strongest instinct: handing back something that sounds like a goal but can't be measured. Ask any model for key results and you'll get "increase user engagement" or "improve onboarding." Those aren't key results. They're aspirations with no number attached, and a quarter later nobody can say whether they happened.
The prompts floating around are either a single line ("act as an OKR coach") or a narrated walkthrough of someone chatting with ChatGPT until it produces something usable. Neither locks an output contract. That's the whole problem, because OKRs live or die on measurability, and a loose prompt produces loose metrics every time.
This covers a prompt that forces a baseline and a target on every key result, flags vanity metrics, and balances leading against lagging indicators.
What you can do with this prompt
- Turn a strategic theme into a structured objective with three measurable key results
- Force a baseline and a target onto every key result, no exceptions
- Catch vanity metrics before they reach the planning doc
- Balance leading indicators against lagging ones so the set is steerable
- Reject output-style "ship feature X" key results automatically
- Draft a full quarter of team OKRs in the same shape
Why models default to unmeasurable key results
A model trained on the open web has read a million blog posts that use "engagement" and "growth" as if they were metrics. So that's what it returns. The fix isn't a better request. It's a contract that rejects anything without a number, a baseline, and a target.
"Increase activation from 31% to 45% by end of Q3" passes. "Improve activation" doesn't. The prompt should refuse the second and ask for the data behind the first.
The single rule that fixes most AI-drafted OKRs: every key result needs a current number and a target number. "Grow MRR to $50k" is incomplete without knowing you're at $32k today, because the target only means something against the baseline. Bake "baseline → target" into the contract and the model stops producing directional fluff. If it doesn't know the baseline, it should mark it as an input you owe it, not invent one.
Anatomy of the prompt
Variables
{{theme}} — the strategic goal or focus area
{{baselines}} — current numbers you already track
{{horizon}} — the period (e.g. "Q3", "this half")
Prompt
Role: an OKR coach who refuses vanity metrics.
Task: turn {{theme}} into one objective and exactly three key
results for {{horizon}}. Every KR needs a baseline and a target.
Use {{baselines}} where available; mark missing baselines as
"NEED BASELINE", never invent one. Include at least one leading
and one lagging metric. Reject any "ship X" output metric.
Output contract
Objective: <one qualitative, time-bound sentence>
KR1 (lagging): <metric> from <baseline> to <target>
KR2 (leading): <metric> from <baseline> to <target>
KR3: <metric> from <baseline> to <target>
Flagged: <any KR that's borderline vanity, and why>
The "reject any ship X output metric" line matters more than it looks. Teams constantly write "launch the new dashboard" as a key result. That's a task, not an outcome. The prompt should push back and ask what the dashboard is supposed to move.
How models behave on this specific job
Behavior on drafting measurable OKRs, not on strategy in general:
| Behavior on OKR drafting | Claude | ChatGPT (GPT-4o) |
|---|---|---|
| Respects the baseline-to-target contract | Steady once the rule is set | Steady, occasionally drops the baseline on KR3 |
| Distinguishes leading from lagging metrics | Reliable | Reliable, sometimes mislabels a lagging metric as leading |
| Invents baselines when none given | Marks NEED BASELINE with the rule | Marks NEED BASELINE with the rule; invents more without it |
| Catches its own vanity metrics | Flags them when asked | Flags fewer unless the example is in the prompt |
The pattern holds across both: with the contract, the drafts are useful; without it, both models produce confident vanity metrics. GPT-4o needs a worked example of a flagged metric in the prompt to flag reliably, while Claude does it from the instruction alone. And both will fabricate a baseline if you let them, so the NEED BASELINE convention isn't optional.
1. Name the theme
{{theme}} is the focus, not the metric. "Make activation the growth engine" beats "increase activation," because the objective should be qualitative and the key results carry the numbers.
2. Feed your real baselines
{{baselines}} is where the draft gets honest. Paste the numbers you already track. The model uses them and flags the gaps, instead of inventing a plausible 28%.
3. Set the horizon
{{horizon}} keeps the targets bounded. An OKR with no end date isn't an OKR. A quarter or a half is the usual frame.
4. Run and challenge the flagged list
Read the "flagged" section first. Those are the key results the model itself suspects are vanity. If it flagged nothing, your theme was probably too output-shaped, so sharpen it and rerun.
5. Fill the NEED BASELINE gaps
Each NEED BASELINE is a number you owe the set before it's real. Filling them is quick. Skipping them leaves you with targets floating in space.
Variables you'll set
| Variable | Required | What it is |
|---|---|---|
{{theme}} | Yes | The strategic goal or focus area for the period |
{{baselines}} | No | Current numbers you already track; gaps get flagged |
{{horizon}} | Yes | The time period the OKR covers |
An opinion worth holding
AI doesn't write your OKRs, and any post claiming it does is selling something. What it does well is enforce the discipline you'd skip under deadline: it won't let "improve retention" through, it forces the baseline, it catches the task masquerading as an outcome. That's worth more than the wording. The wording you'd fix anyway. The discipline is the part teams quietly drop when planning week gets compressed.
So treat the prompt as a structural editor, not an author. You bring the strategy and the numbers. The contract makes sure the result is measurable instead of inspirational. A reusable prompt keeps that contract in one place, so every team's OKRs land in the same trackable shape instead of each manager freelancing the format.
Getting started
- Copy the skeleton into your model.
- Set
{{theme}}qualitative and paste real numbers into{{baselines}}. - Set the
{{horizon}}to your planning period. - Run it, read the flagged list, and fill every NEED BASELINE.
- Challenge any key result you couldn't defend to a skeptical exec.
- For a connected workflow that audits OKRs against your roadmap and execution, use a pack built for alignment, like the Connected Roadmap Alignment Agent Pack.
The Connected Roadmap Alignment Agent Pack keeps the OKRs honest after you draft them: it audits initiatives against strategic objectives, surfaces dependency risks, and publishes a revised recommendation, so the {{theme}} you set in planning doesn't drift from what the team actually ships. It's part of The Complete AI Prompts Bundle, a one-time lifetime license to the full catalog plus future packs, which earns its keep if you run roadmap and requirements jobs every quarter too.
OKRs are the top of a stack that runs down to tickets. The Connected Product Requirements Agent Pack turns an objective into a reviewed PRD with linked tickets, which closes the loop from goal to work. For the layers below, the product roadmap prioritization prompt covers sequencing the bets, and the PRD prompt walkthrough covers the document those bets become.
Browse all strategy and product prompt packs →Common questions
What is an OKR drafting prompt?
Can AI write good OKRs or just reword goals?
Should key results be leading or lagging metrics?
Get the prompt packs this guide is built on
Ready-to-paste prompts with documented variables and usage guides for ChatGPT, Claude, and Gemini. One-time payment, own it forever.
More prompt guides

Gemini vs Claude for Long-Context Code: Window or Accuracy
The honest framing of Gemini vs Claude for long-context code isn't which model is smarter. It's a tradeoff between two different things: how much code you can fit in one prompt, and how often the mode…

A Technical Design Doc Prompt That Holds the RFC Structure
A technical design doc prompt earns its keep when it stops every author from inventing a new doc structure. Context, the options you considered, why you picked one, what breaks, how you roll it out. S…

A Customer-Facing Release Notes Prompt That Drops the Jargon
A customer-facing release notes prompt has one job the engineering changelog never had: talk to someone who doesn't read code. "Refactored the auth middleware" means nothing to a user. "Sign-in is fas…