Foundation

You can't win prompts you haven't named

Why the prompt library — The Brief — is the real foundation of AEO.

By the Sourceworks team

Published January 2026

8 min read

Verification Stack — illustration of interlocking schema layers and entity verification for AI search foundations

Every AEO programme stands or falls on the same thing — and most enterprise teams have never actually built it. The prompt library is The Brief. The 100–200 questions you've explicitly decided your brand belongs inside when buyers turn to AI for guidance. Without it, every downstream metric is hollow, every content investment is a guess, and every dashboard is a polished form of self-deception.

This is the work nobody wants to start with. It's slow, it's editorial, it requires judgement no tool can supply. And it's the single most defensible deliverable in the entire methodology, because every measurement that comes after — citation rate, share of voice, sentiment, source citations — is downstream of which prompts you picked.

Pick badly, and the dashboard tells you nothing. Pick well, and the next twelve months of work have a target.

Why this stage exists at all

AEO has no SERP. That's the bit most teams keep forgetting. In SEO, the keyword universe is fixed, the leaderboard is public, and you can argue about positioning because there's a position to argue about. In AEO, the response varies — by user, by session, by region, by platform. The same prompt asked twice in a row, frankly, can return two different answers. You're not optimising for a slot; you're optimising for inclusion in a generated answer. Which means the question “are we winning?” has no meaning at all until you've defined the prompts.

The Brief is two things at once. It's the list of conversations the brand belongs inside — the editorial decision about which generated answers you intend to be cited in. And it's the measurement frame — the fixed set of inputs you run repeatedly, month over month, to detect movement. Without it, you're producing content into a void and reporting on metrics that don't connect to anything specific.

Of course, you can run any prompt through any visibility tool. Otterly, Frase, HubSpot AEO, Peec — they'll all happily test whatever you feed them. But none of them will decide which prompts are worth feeding. That decision is the human work, and it stays the human work. It's also the work that genuinely cannot be outsourced to a generalist marketing manager — not because they're not smart, but because building a prompt library well requires sitting inside the buyer's vocabulary, not the seller's. Most internal teams, weirdly, are the worst-placed people in their own company to do this.

The Brief is what everything else gets pointed at. If it's right, the rest of the system has a target. If it's wrong, nothing in the system can save it.

The playbook

Five moves, in order. Total time for a serious enterprise build is 15–25 hours over a fortnight — though the better operators stretch it to four weeks to let the buyer-language work breathe.

Move 01

Start with the buyer, not the product

Run a proper intake before generating a single prompt. The temptation — especially for in-house teams — is to skip straight to “what should we be cited for?” That's the seller's question. The buyer's question is “how would I phrase this if I were the person trying to solve the problem?” Those two questions, weirdly, rarely produce overlapping language.

A useful intake covers:

ICP definition

Role, seniority, company size, the team they sit in. Get specific — “marketing leaders” isn't an ICP, “VP-level demand gen leaders at 200–500-person B2B SaaS companies” is.

The Moment of Intent

At what point in their process does the buyer actually open ChatGPT? Before they evaluate vendors? During evaluation? To validate a shortlist they've already built? The answer changes which prompts matter.

Top 5–10 competitors

Both direct and adjacent. Adjacent matters more than people think; buyers shop laterally.

Sales objection patterns

What do buyers ask about that the team knows competitors win on? These are gold for decision-stage prompts.

Existing language assets

Sales call recordings (Gong, Chorus), support tickets, lost-deal notes, customer reviews. Real buyer phrasing lives in these. Marketing collateral lives in seller phrasing. Don't confuse the two.

This sounds basic. It is basic. It's also where almost every prompt library we audit went sideways — they were built off a brand brief, not off buyer evidence.

Move 02

Generate wide, then narrow

Go for 150–250 raw prompts in the first pass. Wide first, sharp second. Five sources, in this order:

Direct AI generation

Ask ChatGPT, Claude, and Perplexity directly to suggest the prompts people use when researching the category. In Perplexity specifically, watch the “related questions” suggestions at the bottom of each response — they're algorithmically derived from what users actually ask next, which makes them surprisingly clean signal.

Competitor reverse-engineering

Pull 3–5 competitor FAQs, comparison pages, and top-trafficked blog posts. Each one is an implicit prompt the competitor's team thought was worth answering.

Buyer language sources

Reddit threads in the relevant subreddits, Quora, customer support transcripts, sales call snippets. This is where you find the prompts your buyer would actually type — phrases like “is X actually worth it” or “alternatives to Y for a small team” that no marketer would ever write down on their own.

Journey mapping

Force-generate prompts across all four stages: Awareness (“what is X”), Consideration (“best X for Y”), Decision (“X vs Y”, “X reviews”, “X pricing”), Post-purchase (“how do I do Z in X”, “X alternatives” — the last one is a churn signal worth tracking).

The branded layer

“What is [Brand],” “Is [Brand] legit,” “[Brand] reviews,” “[Brand] vs [top competitor].” These often surface the most embarrassing data — outdated descriptions, mis-attributed products, the occasional hallucinated executive. Worth knowing about now rather than later.

Move 03

Cluster and tag

Now structure the raw list. Every prompt gets columns:

Funnel stage

Awareness / Consideration / Decision / Post-purchase.

Intent type

Informational / Commercial / Navigational / Transactional.

Topic cluster

Group related prompts; typically 3–7 clusters emerge naturally.

Persona

If multiple ICPs.

Priority

1 = monthly tracking, 2 = quarterly, 3 = annual spot check.

Cull to about 100–150 priority-1 prompts. Fewer than 50 and you don't have statistically meaningful data to work with — citation drift will swamp the signal. More than 300 and the operational overhead eats the engagement.

Distribution rule of thumb: 30% awareness, 40% consideration, 25% decision, 5% post-purchase. Adjust by category — enterprise B2B skews heavier to decision-stage; consumer ecommerce skews heavier to awareness.

Move 04

Test for the baseline

This is the manual labour step. There's no shortcut.

Run every priority-1 prompt through at least three platforms — ChatGPT, Perplexity, and Google AI Overviews are non-negotiable. Add Claude if the audience skews technical; add Gemini if the buyer lives in Google Workspace. Use incognito browsing, log the time of day, and take the run seriously — this is the “before” snapshot that justifies the next twelve months of work.

For each prompt × platform combination, capture:

Mention

Whether the brand was mentioned (Y/N).

Position

Position in the response (first / second / third / list-only / not at all).

Sentiment

Positive / neutral / negative / mixed.

Competitors named

Top 3 competitors named in the response.

Source URLs cited

Which third-party domains the AI leaned on for the answer.

Response excerpt

A representative excerpt for the file.

Two-pass it to stay sane: first pass is just Y/N — gives you the “are we on the map?” answer in about 90 minutes for 150 prompts. Second pass is the detailed capture for the prompts that matter most.

Cross-check with the free tools — HubSpot AEO Grader, Ahrefs Brand Radar, Semrush's AI Search Visibility — and screenshot the outputs. They become the visuals in the baseline report and they triangulate the manual numbers, which matters because every tool has its own blind spots.

Move 05

Set the targets in writing

Last move, and the one most teams skip: explicitly agree, in writing, what success looks like at 90 days, 6 months, and 12 months. Which prompt clusters matter most, ranked. What citation rate target is realistic — enterprise baselines typically come in at 5–15% on priority-1 prompts, and a credible 12-month target is 25–40%. Which engines matter most for this audience, and which ones are being explicitly deprioritised. How citation drift will be handled in monthly reporting — rolling 4-week averages, not raw monthly numbers.

This step is where most engagements quietly fail later. Without written targets, every monthly report becomes a Rorschach test for whatever the loudest stakeholder feels that week — and a programme that's actually working can get cancelled in month four because the wrong stakeholder read a noisy chart on the wrong morning.

What to measure

Five signals come straight off The Brief. None of them mean anything individually — the dashboard is the unit, not any one metric.

Citation rate

What percentage of priority-1 prompts mention the brand at all, by platform. The base metric. The one the CMO will want.

Share of voice within the cluster

When the brand is mentioned, where is it in the response order, and which competitors share the surface. This is the metric that actually moves stakeholders — “we're being mentioned but always after [main competitor]” is a different problem than “we're not being mentioned at all,” and the response is different.

Sentiment

Positive / neutral / negative / mixed when the brand does appear. Negative or factually wrong sentiment is an entity-foundation problem you want to surface now rather than discover at month four.

Source citations

Which third-party domains is the AI grounding its category answers in. This is the most useful diagnostic in the whole report — it tells you exactly where the next 6 months of distribution work needs to happen.

Coverage by engine

How performance distributes across ChatGPT, Perplexity, Claude, Gemini, and Google AI Overviews. Different engines pull from different source mixes; the spread is informative on its own.

Report on rolling averages, never raw monthly numbers. Citation drift — the natural variance in which sources AI systems pull for the same prompt — produces 15–30% month-to-month swings that look like catastrophes in a line graph and like noise in a quarterly view. The team that doesn't understand drift will read every dip as failure. The team that does will absorb it without flinching.

Common failure modes

Five patterns, almost all of them avoidable, almost all of them present in the first audit we run for a new client.

Reading single months as truth

Citation rate dropped from 28% to 19% in October — what's wrong? Honestly, probably nothing. Probably citation drift. The same prompt, asked twice in a row, can return two different answers; over a month, the variance compounds. The instinct is to redirect work in response. The discipline is to wait for the rolling 4-week average to settle and check whether the trend is real. Most aren't. This is the single most counterintuitive thing about AEO measurement, and the teams that learn it in month one survive their first bad month. The teams that don't, don't.

The Seller's Vocabulary Trap

The prompt library was built off the brand brief, the product taxonomy, or the internal category language. Result: it tracks prompts no buyer ever types. The fix is brutal — throw it out and rebuild from sales call transcripts, support tickets, and Reddit threads. Buyer language almost never matches seller language, and the gap between them is where most of the visibility opportunity actually lives.

Under-populated libraries

40, 60, sometimes as few as 25 prompts being tracked. Citation drift turns that into pure noise. You need at least 100 priority-1 prompts to detect signal versus drift in a 90-day window. Anything less and you're reporting feelings.

No priority tiers

Every prompt treated as equal. Most prompts aren't — a decision-stage comparison query is worth twenty awareness queries to the business — and a flat library produces dashboards where movement on the high-value prompts gets washed out by noise on the low-value ones.

Tool worship

The team bought a dashboard, populated it with whatever prompts the tool suggested, and called it a prompt library. It's not. A vendor tool can run prompts; it cannot decide which prompts matter for this specific buyer, in this specific category, at this specific point in their journey. That work stays human. The tool is the instrumentation, not The Brief.

Before you publish more, measure what AI already believes.

Request a Diagnostic →