The Problem Nobody Is Measuring Correctly
Most teams chasing AI search visibility are optimising for the wrong outcome. They want citations — a URL appearing somewhere in a ChatGPT or Perplexity response. That is a reasonable starting point, but it conflates two distinct events: citation selection (the model pulls your URL into a response) and citation absorption (the model actually uses your content to form its answer, not just appending your link as a footnote).
We ran 17 audits across 14 distinct domains since our launch in May 2026 and found an average GEO score of 23 out of 100 — a straight Grade F, extracted from audit data as of 2026-05-26. Not a single domain we audited had passed the citation-absorption threshold on more than two of the four major AI platforms simultaneously. The gap between "we rank on Google" and "we are used by AI engines" is not a content quality problem alone. It is a structural measurement problem: businesses are not tracking the right signals, so they cannot fix the right things.
Two Concepts You Must Separate
Before building a measurement framework, you need a clean conceptual split. These are not synonyms:
| Concept | Definition | What it looks like in a response |
|---|---|---|
| Citation selection | The model includes your URL in its source list | Your domain appears in Perplexity's "Sources" panel |
| Citation absorption | The model draws on your content to construct its answer | A claim, figure, or framing in the answer traces directly to your page |
| Citation displacement | A competitor's content replaces yours between model updates | Your URL disappears from responses it previously populated |
| Citation latency | Delay between content publication and first AI appearance | New pages taking 4–10 weeks to surface in Gemini responses |
| Cross-platform parity | Your content performs consistently across ChatGPT, Claude, Gemini, Perplexity | Most sites are cited on one platform but absent from three others |
| Structured recall | AI can reproduce your facts, definitions, or steps verbatim | Model answers a definition question using your exact framing |
The academic groundwork for this distinction exists. Aggarwal et al. published the paper that formally named the field: "we introduce Generative Engine Optimization (GEO), the first novel paradigm to aid content creators in improving their content visibility in generative engine responses" — Pranjal Aggarwal, lead author, GEO: Generative Engine Optimization, KDD 2024, arXiv:2311.09735. That paper established that citation frequency alone is an insufficient proxy for visibility — the position and weight of a citation within a generated answer matters just as much as its presence.
A more recent preprint extends this thinking: arXiv:2604.25707v2 examines how retrieval-augmented generation systems select and weight sources, showing that structural signals — schema markup, authoritative internal linking, and factual density — predict absorption more reliably than domain authority metrics inherited from traditional SEO.
The Six-Signal Measurement Framework
Here is the framework we apply in every Cyberg7 AEO audit. Each signal is binary or numeric so you can track it in a spreadsheet before you invest in tooling.
Signal 1 — Citation Frequency Score (CFS)
Run a fixed prompt set (minimum 20 queries covering your target topics) across ChatGPT, Claude, Gemini, and Perplexity. Count how many responses include your domain in the source list. Express as a percentage per platform. A site scoring above 40% on any single platform has cleared the baseline citation-selection threshold.
Signal 2 — Answer Alignment Rate (AAR)
For each response where you were cited, read the generated answer and score whether a claim, definition, or data point in that answer is traceable to your page content (1 = yes, 0 = no). AAR = absorbed citations ÷ total citations. Most Grade F sites we audit have an AAR below 0.15, meaning the model cites them but ignores what they actually say.
Signal 3 — Cross-Platform Parity Index (CPPI)
Divide the number of platforms where CFS > 20% by 4 (the four major platforms). A CPPI of 1.0 means you appear consistently everywhere. A CPPI of 0.25 — one platform only — is the most common pattern we see among B2B SaaS sites in Singapore and Southeast Asia. Google's own documentation on its AI-powered search features notes that source selection varies by query type, which partly explains why cross-platform parity is hard to achieve without deliberate structural optimisation.
Signal 4 — Structured Recall Score (SRS)
Ask the AI engine a direct definitional or factual question your page is designed to answer. Paste the response back to your page and check whether the AI reproduced your framing, your specific numbers, or your named framework. Score 1–3: 0 = no match, 1 = partial overlap, 2 = clear paraphrase, 3 = verbatim or near-verbatim use. This is the highest-fidelity signal for absorption.
Signal 5 — Citation Displacement Velocity (CDV)
Run the same prompt set monthly. Track when your domain drops out of responses it previously populated. High displacement velocity (losing more than 20% of citations month-over-month) signals that a competitor has published content with better structural signals and the model has reweighted its retrieval. Search Engine Journal's coverage of generative search trends consistently highlights that model updates — not just content updates — can trigger sudden citation displacement.
Signal 6 — Latency to First Appearance (LFA)
Publish a new page with a clear, dateable claim (a specific statistic, a named framework, a singular definition). Start running your prompt set the day of publication. Log the first date the page appears in any AI response. Across our audit cohort, LFA ranges from 11 days (for pages with strong schema markup and inbound links) to never (for pages with no structured data and thin internal linking).
How to Build This Into an Operational Cadence
You do not need an enterprise analytics stack to start. Here is the minimum viable measurement cadence:
- Define a fixed prompt set of 20–30 queries covering the topics your business needs to own. Write them as a user would ask them — conversational, specific, and without your brand name in the query.
- Run prompts manually across all four platforms on a fixed cadence (weekly for new content, monthly for stable pages). Paste results into a shared spreadsheet with columns for date, platform, query, cited (Y/N), absorbed (Y/N), and SRS score.
- Calculate CFS, AAR, and CPPI at the end of each month. These three numbers give you a one-page dashboard of where you stand.
- Set a CDV alert threshold. If any topic cluster loses more than 15% of its CFS in a single month, flag it for content review — the signal is that a competitor or a model update has pushed your content down in retrieval weighting.
- Add schema markup to every content page before its first prompt-set run. The arXiv:2604.25707v2 preprint cited above and our own audit data both point to structured data as the single highest-leverage structural intervention for improving LFA and AAR simultaneously.
- Review your GEO fundamentals against a published framework. Our explainer at What is GEO? Generative Engine Optimization Explained covers the foundational content signals — authoritative tone, factual density, direct answers — that feed into absorption scores. Use it as a checklist alongside this measurement framework.
- Rerun a full audit every quarter. Measurement without a reset point becomes noise. A quarterly full audit (the same 20–30 prompts, all platforms, all six signals scored) gives you the trend line that matters to a founder or CMO.
Why Grade F Is the Baseline and How to Move Off It
The average score of 23/100 across our first 17 audits is not a fluke of sample selection. We audited well-run B2B sites — companies with existing Google rankings, published thought leadership, and active content calendars. The problem is not effort; it is that those sites were built and measured entirely inside a Google-centric SEO paradigm. AI engines retrieve and weight content differently. They favour direct, structured answers over SEO-optimised prose; they favour pages that demonstrate expertise through specific claims over pages that signal expertise through keyword density.
Citation absorption — Signal 2 in this framework — is the metric that most directly predicts whether AI search will generate a lead for you or for a competitor. A cited page that the model ignores when constructing its answer contributes zero to your pipeline. A page with an AAR above 0.5 is actively doing sales work every time a prospect asks an AI engine a question in your category.
The measurement framework above is designed to make that difference visible, trackable, and improvable. Start with the prompt set. Score it this week. The number you get will tell you exactly how much structural work is ahead of you.
Related Reading
Run Your Own Audit
If you want a scored baseline across all six signals rather than building the prompt set yourself, run your free AEO audit at aivisibility.cyberg7.com.sg/audit — you will get a platform-by-platform CFS breakdown and a prioritised fix list within 24 hours.
Newsletter
Get the next post the day it ships
One short field note per week — frameworks, audit data, and tactical guides for AI search visibility. Unsubscribe anytime.
We only email when we ship a new post. No drip campaigns, no upsell sequences.