Measurement

Measuring AEO without making things up

Why The Proof Layer — the five metrics, the rolling averages, the reporting cadence, and the attribution model that survives a board meeting — is what stops AEO from becoming the next line item finance cuts.

By the Sourceworks team

Published January 2026

10 min read

Verification Stack — illustration of interlocking schema layers and entity verification for AI search foundations

AEO measurement is the stage where most enterprise programmes quietly die. The dashboard looks busy, monthly numbers swing 10–15% in either direction for reasons nobody can explain, the CMO can't tell a CFO what's actually moving, and by the third quarterly review the budget conversation has already turned — finance has decided the line item isn't defensible, and the work goes with it, regardless of how good the work was.

Most enterprises don't have one. Most consultancies don't sell one. Most teams running AEO programmes are producing data and calling it measurement.

The Proof Layer is what stands between that quiet death and a programme that survives. It's the framework, the cadence, and the attribution model that turn five inherently noisy metrics into a story finance can sign off on at a budget meeting — and discover the difference at exactly the worst possible moment if you don't have one in place.

Why this stage exists at all

AEO measurement is harder than SEO measurement for a structural reason. There's no SERP. There's no shared, public leaderboard every team can point to and argue about. AI responses vary across users, sessions, regions, and platforms — the same prompt asked twice in a row can return two different answers. SISTRIX's 2026 citation drift study found 56–74% of sources cited by AI rotate week-over-week. The volatility isn't a bug in the measurement; it's a property of the medium. A measurement framework that doesn't account for it produces noise so loud it looks like failure every other month.

The default response — the response we see in almost every measurement review we run — is to buy a tool, log in to its dashboard, and call that measurement. It isn't. A tool dashboard is just data. Measurement is the framework around the data: which prompts are being tracked and why, what “good” looks like for this specific business, what cadence absorbs the noise without losing the signal, which tools are being triangulated across, what the rolling averages actually say once the drift is stripped out, and — frankly the most important part — how the report reads to three different audiences at once: a CMO who wants the trend, a CFO who wants the line, and a board who wants the story.

The Proof Layer is also where AEO becomes defensible at the line-item level. The work in Foundation, Authority, and Distribution is the substance — but Measurement is what proves the substance is moving. Without it, AEO sits next to “brand campaigns” in the marketing budget as another initiative the team can't quite justify when pressed. With it, AEO sits next to “infrastructure” — and infrastructure is a different conversation entirely.

Brand campaigns or infrastructure. The Proof Layer decides which conversation finance has about you.

The playbook

Five moves, in order. The first four set up the system. The fifth is the work that decides whether the engagement renews at the twelve-month mark.

Move 01

Build the framework before any execution work

The framework is the document everything else runs against. It defines what's being measured, why, how often, against what targets, and what the report looks like when it lands. It gets built once at the start of the engagement, signed off in writing by the client's leadership, and refreshed annually. The five-metric set has now standardised across the industry — HubSpot's vocabulary, since the AEO tool launched in April 2026, is becoming the de facto default. The five metrics:

Citation rate

Percentage of priority prompts where the brand appears in AI responses. The base metric. The one the CMO will want first. Calculated against The Brief — the prompt library from Foundation — so the denominator is always the prompts the client has explicitly decided to compete inside. Pick the prompts badly and citation rate measures nothing useful.

Share of voice

Brand mentions divided by total brand mentions across the tracked set, per prompt cluster. The metric that actually moves stakeholders, because “we're being mentioned but always after [main competitor]” is a different conversation than “we're not being mentioned at all,” and the response is different.

Sentiment

Tone of mention when the brand does appear. Negative or factually-wrong sentiment is usually an entity-foundation problem — broken Verification Stack, missing Trust Chain, outdated descriptions on third-party sources. Surface it now rather than discover it at month four.

Coverage by engine

Citation rate broken out across ChatGPT, Perplexity, Claude, Gemini, Copilot, and Google AI Overviews. Different engines pull from different source mixes — the spread is informative on its own and tells you where the next quarter's work needs to land.

AI referral traffic

Sessions arriving from AI sources, tagged in GA4. Small in absolute terms, structurally undercounted (more on that in Move 5), but the only direct link to the analytics infrastructure the rest of marketing already operates on.

Setting targets is the work most teams skip — and the work that protects the engagement when the inevitable bad month arrives. Targets get written down explicitly, with three horizons: 90-day (modest, mostly noise), 6-month (where the trend becomes interpretable), 12-month (where the case for next year's budget actually gets made). Sample target language: “Move citation rate from 12% to 25% across priority-1 prompts over six months. Move share of voice from 18% to 30% relative to top three competitors over nine months.” Without targets, every report is a Rorschach test. With them, the report has a structure the board reads in two minutes.

Move 02

Triangulate the tools — never trust one

Every AI visibility tool has blind spots and incentives. Some mark too many mentions as positive. Some miss specific platforms entirely. Some have AI modules bolted onto legacy SEO products that are still maturing. The honest framing is that no single tool is reliable enough to make decisions from — which is why the rule is always two, sometimes three, in parallel.

The minimum stack:

One free tool for baseline

SISTRIX for AI's free tier is the strongest here, particularly for the citation drift dimension. Used for quarterly cross-checks rather than monthly reporting.

One paid tool for monthly tracking

Otterly at the entry tier, HubSpot AEO if the client is already on Marketing Hub, Peec if the analytics depth matters, Frase for content-side integration. Pick one and commit; switching mid-engagement breaks the trend baseline.

Bing Webmaster Tools + AI Performance dashboard

The only first-party AI citation data currently available, shipped by Microsoft in February 2026. Free. Mandatory.

GA4 segmented for AI referral sources

Custom segments, source/medium overrides for AI tools, and a clean tagging convention applied across the analytics infrastructure the rest of marketing already operates on. The bridge between AEO measurement and everything else the team reports on.

The discipline that matters: when two tools disagree, that's signal, not failure. The reconciliation work — figuring out why Tool A shows the brand cited in 32% of priority prompts while Tool B shows 24% — is where the actual judgement lives. The consultancy that gives the client a single number from a single dashboard is, frankly, selling the appearance of measurement rather than measurement itself. One operational warning: don't over-tool. Three is the ceiling; even at three, data reconciliation eats hours that should be going into interpretation. A single thoughtful tool plus the Bing dashboard plus GA4 is a stronger framework than five tools running in parallel that nobody has time to read.

Move 03

Engineer for citation drift on purpose

Citation drift is the central methodological challenge of the entire Proof Layer. It is not a problem to solve. It is a property of the medium to absorb.

The SISTRIX 2026 study established the baseline numbers: 56–74% of sources cited by AI rotate week-over-week, depending on platform. Google AI Overviews is the most stable — over half of tracked prompts show zero source change ever, which is its own kind of weirdly informative finding. Google AI Mode shows a core-plus-carousel pattern, where a small set of citations stay locked in and the surrounding references rotate. ChatGPT is the most volatile, with near-total weekly fluctuation in many categories.

Three platform architectures, three measurement behaviours, one principle: read trends, not snapshots. The operational implications, in order:

Rolling averages, not monthly numbers

Every chart in every report shows the most recent month alongside a 12-week rolling average. Visually, you can see exactly which line is signal and which line is noise. The monthly numbers swing. The rolling average tells the truth.

Per-platform reporting

Don't aggregate ChatGPT's volatility with AIO's stability and call the average meaningful. The story is in the platform-by-platform read.

Domain-level over URL-level

A specific URL being cited matters less than the domain being consistently cited. Track domain presence; let URL-level be the diagnostic for which content is doing the work.

The “How to Read These Reports” document

A one-page deliverable that ships with the framework, summarising drift, explaining the rolling-average view, and setting expectations for the first three reports explicitly. Month 1: this is your baseline, no conclusions. Month 2: pipeline activity, mostly still noise. Month 3: first inflection points, trend unstable. Month 4 onward: trend is interpretable. Giving the client this document in writing, on day one, is the single highest-leverage piece of paper in the engagement.

The team that absorbs drift without flinching renews the budget at year-end. The team that doesn't, doesn't.

Move 04

Run the three-tier cadence

Monthly absorbs the noise. Quarterly tells the story. Annual makes the case. Each has a different job; running them as the same job is one of the most common failure modes in the entire stage.

Three tiers, three jobs:

Monthly — operational

Eight to twelve pages, delivered by the 5th of the following month, with a 30-minute walkthrough call within the first week. Executive summary, metrics dashboard, trend visualisations with annotations, engine-by-engine breakdown, prompt-cluster performance, competitive intelligence, sentiment and source citation analysis, AI-influenced business outcomes, this month's activity, next month's plan. The point isn't to interpret — it's to maintain rhythm and surface anomalies early. Most months read “trend continues; no major intervention required.” That's a feature, not a problem.

Quarterly — editorial

Twelve weeks of data analysed as a single trend. A citation source map — which third-party domains AI is grounding answers in for the category, compared to last quarter — which feeds The Mention Graph and The Citation Surface directly. A sentiment deep dive, which feeds The Trust Chain. Competitive positioning. A prompt library audit — what's no longer relevant, what should be added. And strategy adjustments for the next quarter, grounded in what the data is actually saying. Fifteen to twenty-five pages. This is the report that gets forwarded.

Annual — the budget review

Twelve months of trend analysis. A stage-by-stage retrospective across all eight build moves of the methodology — what Foundation produced, where Authority compounded, which Distribution channels delivered, what the measurement infrastructure itself surfaced. Industry context. Strategy for the year ahead. And a clear answer to the question the CFO is going to ask before anyone else gets to: what did this twelve months actually buy us? The annual has to defend itself in the room. Bringing a thin annual to a budget meeting is how you lose the account.

Three cadences, three different muscles. The team that runs all three keeps the engagement. The team that flattens them into a single monthly grind — because monthly is what the client asks for — loses the strategic case for the work at the budget review.

Move 05

Build the attribution model honestly

The hardest measurement question every client eventually asks: does this work drive business outcomes? Most consultancies don't have a good answer. The honest answer is that AEO attribution is structurally partial — and naming the limits up front is what makes the rest of the framework credible. Three things are true at once: AI referral traffic is genuinely small in absolute terms; the gap between AI influence and AI attribution is wide (someone reads a ChatGPT recommendation, opens a new tab, types the brand into Google, lands via “organic” — and analytics credits the wrong channel); and B2B buying cycles compound the gap further, with AI surfaces invisible by the time the deal closes months later. The Proof Layer doesn't solve this. It names it, and builds the partial-attribution model that holds up anyway. Three components:

Multi-touch attribution in GA4

Tag AI referral as a touchpoint — first, mid, or last touch. Identify conversions where AI was any touchpoint. Compare conversion rates for AI-touched users versus non-AI-touched users. This catches the visible portion of AI influence.

Branded search lift correlation

Track branded query volume in Google Search Console over time and correlate against AI visibility work. A spike in branded searches after a major byline lands is a strong directional signal — it catches the “saw an AI recommendation, then Googled the brand” pattern that direct attribution misses.

CRM-tagged AI influence

For B2B specifically: add a “How did you hear about us?” field and train sales to capture mentions of ChatGPT, Claude, Perplexity, AI Overviews. Tag deals with AI influence when prospects mention researching via AI. Partial, but honest.

The under-claim is what makes the rest of the number credible. The consultancy that claims definitive AI attribution is selling a model that doesn't exist.

What “good” looks like

The framework holds up at a board meeting if five things are in place. If they are, the CFO can read the report in three minutes. If they're not, the work survives until the next budget review and then doesn't.

Five metrics, documented baselines, written targets

All five metrics tracked against documented baselines and explicit 90/6/12-month targets, signed off by leadership in writing before any execution starts.

Two tools in parallel, Bing dashboard included

Two tools running in parallel (three at the outside), with the Bing AI Performance dashboard included as standard. Reconciliation work happens openly in the quarterly.

Rolling averages lead the narrative

Reporting on rolling 12-week averages with the monthly numbers shown alongside but not leading. The trend tells the story; the month is the diagnostic.

Attribution model that names its limits

Partial, multi-touch, AI-influenced pipeline as the headline business-outcome number — never a single deterministic claim. The caveat about structural attribution gaps is written into every report.

Leadership success-criteria conversation

Before any execution work begins, the framework gets walked through with the CMO, the CFO if accessible, and ideally one board observer. The point isn't to negotiate the metrics — it's to make sure the leadership team has read them, agreed them, and committed to interpreting them through the same lens. That conversation, in writing, is what makes everything that follows defensible.

Common failure modes

Five patterns, present in almost every measurement review we run, all the difference between a programme finance renews and one finance quietly cuts.

The Dashboard Mistake

The team buys Otterly, Peec, or HubSpot AEO, logs into the dashboard, and treats that as measurement. It isn't. A dashboard with no defined success criteria, no rolling averages, no documented baseline, no competitive view, and no agreed targets is just data — and data isn't a story a CFO can read. The fix is to build the framework first, before any tool is selected. The tool serves the framework, never the other way around.

The Single-Number Trap

The CMO asks for a single headline metric to take into the board meeting. The temptation is to give them one — “we're up 14%” — and the cost is enormous. Citation rate without share of voice is one-dimensional. Share of voice without sentiment misses the entity problem. Sentiment without coverage by engine flatters volatile platforms. The dashboard is the unit. The single-number version is what produces month-three panic when the single number swings. The fix is to train the CMO to read the dashboard the same way the measurement team does — five metrics, rolling averages, against documented targets. Resist the headline-number request until the leadership team has read three reports together.

Citation Drift Panic

Month two, citation rate dropped from 28% to 19%. The client asks what's gone wrong. Honestly, probably nothing — probably drift. The instinct is to redirect work, change strategy, ship new content faster. The discipline is to wait for the 12-week rolling average to settle and check whether the trend is real. Most aren't. The fix is the “How to Read These Reports” document delivered in month one, the rolling-average view in every monthly report, and the explicit expectation-setting language baked into the first three reports. Teams that get this right survive their first bad month. Teams that don't, don't.

The Fake Ranking Report

The client wants something that looks like an SEO ranking report — “we ranked #3 for [prompt].” The honest answer is that ranking, in the SEO sense, doesn't exist in AEO. The consultant who manufactures one to satisfy the question is undermining the rigour of every other metric in the report. The fix is to reframe the question — from “what's our ranking?” to “what's our citation occupancy and share of voice within this prompt cluster, and how is it trending?” — and accept the short-term friction of teaching the new mental model.

Attribution Overreach

The report claims AI drove $X in pipeline this quarter. The CFO asks how. The answer turns out to involve assumptions the consultancy can't defend, and trust in every other number in the report goes with it. The fix is the opposite — under-claim, name the limits, and let the AI-influenced pipeline metric speak as a directional estimate rather than a deterministic figure. “We can attribute roughly $X with confidence; the true number is likely higher because of structural attribution gaps we'll be transparent about.” That sentence, delivered consistently, protects the entire framework.

Where this connects

The Proof Layer is Move 9 in the methodology — the final move, and the one that decides whether the previous eight earn their place in next year's budget. It's also, weirdly, the move that loops the system back to the start. Every quarter, the prompt library refresh feeds back into The Brief. Every sentiment dip surfaces an entity gap that feeds back into The Verification Stack. Every source citation map points at where The Mention Graph and The Citation Surface need to push next. Every coverage-by-engine read tells The Indexing Layer and The Agent Layer where the next quarter's work has to land. The methodology isn't a list — it's a loop, and The Proof Layer is the joint that closes it.

Build The Proof Layer. The work defends itself.

Before you publish more, measure what AI already believes.

Request a Diagnostic →