What is AEO? How to Get ChatGPT, Perplexity & AI Search Engines to Cite Your Website — 2026 Guide

Google impressions look fine. Traffic is holding. But when a prospect asks ChatGPT "best [your category] tools in Singapore," your brand doesn't appear. That's

Cyberg7 AEO team·AI visibility editorial
·6 min read

Your site is invisible to AI — and you don't know it yet

Google impressions look fine. Traffic is holding. But when a prospect asks ChatGPT "best [your category] tools in Singapore," your brand doesn't appear. That's not a ranking problem — it's an Answer Engine Optimization (AEO) problem.

AEO is the practice of structuring your website so that large language model (LLM) systems — ChatGPT, Perplexity, Claude, Gemini — surface your content as a cited source when they answer user questions. It is distinct from classic SEO, which optimises for a list of blue links. AI engines synthesise an answer and optionally cite the sources they drew from. If your content wasn't crawled, wasn't structured, or wasn't authoritative enough to survive the model's relevance filter, you simply aren't in the answer.

We ran 17 audits across 14 distinct domains since our launch in May 2026 (data extracted 2026-05-26). The average AI visibility score across those sites was 23 out of 100 — a Grade F. These weren't obscure companies: they included B2B SaaS products and regional marketing agencies that rank on page one of Google. Being findable by a crawler and being citable by an LLM are, it turns out, very different things.


The AEO framework: four layers that determine whether you get cited

Think of AI citation as a funnel. Each layer is a prerequisite for the next.

LayerWhat it controlsFailure mode
1. Crawl accessCan GPTBot / PerplexityBot reach your pages?Blocked in robots.txt or behind auth walls
2. Content structureDoes the page answer a specific question, clearly and early?Walls of marketing copy with no direct answer
3. Authority signalsDoes the page carry enough trust signals for the model to cite it?No author, no date, no external references
4. Machine-readable metadataDoes the site have an llms.txt, schema markup, and clean headings?Model can't parse what the page is about

All four layers must pass. A site with great content but a blocked crawler gets zero citations. A site that's crawlable but structured like a brochure gets skimmed and ignored.


What the research says — and what our audits confirm

The academic framing for this problem comes from a 2023 paper titled GEO: Generative Engine Optimization, accepted at KDD 2024 and published on arXiv as arXiv:2311.09735. Lead author Pranjal Aggarwal and co-authors describe the problem precisely: "we introduce Generative Engine Optimization (GEO), the first novel paradigm to aid content creators in improving their content visibility in generative engine responses." That paper is the first peer-reviewed definition of the field, and it matters because it establishes that AI citation is a measurable, optimisable variable — not luck.

On the practical side, a 2026 practitioner guide published on dev.to by PPCvote breaks down AEO into five operational categories: structured content, technical crawlability, authority building, schema markup, and entity clarity. The guide notes that Perplexity in particular rewards pages that answer questions in the first 100 words — a pattern we see confirmed in our own audit data, where pages with a direct answer in the opening paragraph score an average of 18 points higher on our citation-readiness rubric than pages that bury the answer in paragraph four or later.

UltraLab's AEO guide at ultralab.tw adds a useful framing for B2B marketers: AEO is not about replacing SEO but about extending it. The underlying content quality requirements overlap heavily — original research, clear authorship, factual accuracy — but AEO adds a layer of machine-readability that classic SEO never demanded.

On the technical side, OpenAI publishes GPTBot's crawl documentation openly. It tells you exactly which user-agent string to allow in robots.txt (GPTBot), and it confirms that pages blocked to GPTBot are not indexed into ChatGPT's retrieval systems. In every one of our 17 audits, we check robots.txt as step one. Four of the 14 domains we audited had inadvertently blocked AI crawlers — either with a blanket disallow or with a misconfigured wildcard rule.

The fifth technical layer — and the one most sites haven't heard of yet — is llms.txt. Our post LLMs.txt Explained: Write One That Moves Your Score covers this in detail: it's a plain-text file at yourdomain.com/llms.txt that tells AI systems which pages matter, in what order, and what your site is for. It's the robots.txt equivalent for the LLM era, and writing one is a one-hour task that directly improves how models interpret your site hierarchy.


Five steps you can act on today

These are ordered by impact-to-effort ratio, not alphabetically.

  1. Unblock AI crawlers in robots.txt. Check your robots.txt for Disallow: / rules or wildcard blocks. Add explicit Allow directives for GPTBot, PerplexityBot, ClaudeBot, and Googlebot-Extended. Takes 10 minutes. Skipping this step makes every other fix irrelevant.

  2. Write a direct answer in your first 100 words. For every page targeting a question-style query, put the answer before the context. "Callbox.com.sg is a B2B lead generation platform serving Southeast Asian markets" is better than three paragraphs of background before the product description. AI engines extract early text first.

  3. Add an llms.txt file. A machine-readable index of your most important pages gives AI systems a clean map of your site. List your cornerstone pages, your product pages, and your most-cited blog posts. Plain text, Markdown-friendly. See our full guide linked above.

  4. Add FAQ and HowTo schema to your key pages. JSON-LD structured data doesn't guarantee citation, but it dramatically increases the probability that a model correctly interprets your content type. Use Google's Rich Results Test to validate before deploying.

  5. Build your entity footprint. AI models cite sources that appear authoritative across multiple surfaces. This means: consistent NAP (Name, Address, Phone) across directories, a Google Business Profile, author pages with real credentials, and inbound links from domains the model already knows. For Singapore-based B2B companies, getting listed on G2, Clutch, and CNA's tech coverage pages are all high-signal moves.



Run your own AI visibility audit

Not sure where your site sits on the crawl-to-citation funnel? Our free audit at aivisibility.cyberg7.com.sg/audit scores your domain across all four layers in under two minutes and tells you exactly which fixes to prioritise first.

Newsletter

Get the next post the day it ships

One short field note per week — frameworks, audit data, and tactical guides for AI search visibility. Unsubscribe anytime.

We only email when we ship a new post. No drip campaigns, no upsell sequences.

Chat with us