A customer types your category into ChatGPT. Best bookkeeping firm for e-commerce. Best plumber in Austin. Whatever it is you sell.
It comes back with three names and one line on each.
You are not one of them.
You did not lose that deal. You were never in the room, and nothing in your analytics will ever tell you it happened. There is no impression, no bounce, no page two. Just an answer that went to someone else.
Short answer: in a scan of 1,205 live pages, 35.7% returned no usable text to an AI agent at all, and almost none of it correlated with SEO health. The checklist below is the one we run. It takes about twenty minutes and needs nothing but a browser.
Both pages look finished. Only one of them exists to an AI agent.
The difference is not quality, length or design. It is where the content gets assembled. Page A builds itself in the browser, so a crawler that does not execute JavaScript receives navigation and nothing else. Page B ships the same words in the HTML.
Eight words against 488. That is the whole problem in one image, and it is the failure nothing in your analytics reports.
Open the page, hit Ctrl+U, and search for a sentence you can see on screen. If it isn't in the source, an agent can't read it either.
Disable JS and reload. If the page goes blank, many crawlers see the same blank page.
Check robots.txt for the AI crawlers specifically: GPTBot, ClaudeBot, PerplexityBot, Google-Extended. Blocking them is usually accidental, inherited from a template.
If the unique text is under roughly 150 words, there is nothing worth extracting. Half a service page of navigation and a contact form counts as empty.
77.1% of the pages we scanned had no date anywhere on the content. Undated reads as untrustworthy, especially in fast-moving categories.
Only 21.2% of pages had a visible author. "Admin" and "The Team" do not count. An assistant needs to know who is making a claim before it repeats it.
Same legal name, same address, same handles everywhere. If your name renders three different ways across the web, an agent can't merge you into a single thing it can recommend.
This is the ten-second test. Open your own service page and read them. Do they answer the buyer's question, or describe your journey since 1998? An assistant cannot quote the second one.
44.2% of AI citations come from the first 30% of a page (Kevin Indig, 18,012 verified ChatGPT citations). Put the extractable answer there, not in paragraph nine.
62.8% of pages we scanned had none. FAQ blocks are the one exception to front-loading, because each question and answer is self-contained, so they get cited even at the bottom of a long page.
Reviews, directories, community discussion, round-ups. Assistants recommend businesses they see named by someone other than the business.
Each model has a different source diet, and the placement table further down maps it. A decent page on a platform the model already reads beats an excellent page where it never looks.
This is the loop we run monthly. It takes ten minutes per query and it is the only honest way to know.
Not your brand name, the question a customer types before hiring. "Best plumber in [city]", "best local rank tracker", "who should I hire for X".
Separately. They do not agree with each other.
Mentioned in the text is one thing. Cited as a source with a link is another, and worth more.
Pull the source URLs from the footnotes or sidebar.
This is the step that surprises people. Many cited sources barely rank in Google. Ranking well and getting cited are not the same channel.
Add or update a listicle in your category, build mentions on that model's preferred platforms, and sharpen structure and markup on the page you want cited.
Repeat monthly against the same queries. Changing the queries each time makes the result unmeasurable.
Verified by hand across our own tracked queries. This is observation, not a controlled study, so treat it as a starting map rather than statistics.
| Assistant | Leans on | So publish |
|---|---|---|
| ChatGPT | Wikipedia, news agencies, government sites, long reviews and listicles | own-site listicles under buyer-intent queries; encyclopedic and editorial placements |
| Claude | reviews and local sites, plus a broad mix of blogs and guides | directories and review platforms; detailed service and neighbourhood pages |
| Perplexity, Gemini | YouTube, Reddit, large editorial sites | YouTube with real descriptions and timestamps; genuine Reddit participation |
One counter-intuitive rule for listicles: include your competitors. Assistants reward genuine round-ups and discount self-promotion. A comparison page that names rivals outperforms a page that only sells you.
How this squares with the published numbers. Otterly's 2026 analysis of more than a million citations reports a different cut of the same territory: news and media account for 20 to 30% of citations depending on platform, community forums for 5.9 to 16.9%, and brand sites get mentioned often but linked rarely.
Those figures do not contradict the table above. They answer a different question. Otterly measures what share of all citations each source type earns across a broad query set. The table records which source types actually turned up against a specific set of buyer-intent queries in service categories. A source type can be a small slice of the internet-wide total and still be the one that decides your category.
Where the two agree is the part that matters here: both put community and review platforms well ahead of brand-owned pages as citation sources. That is the whole argument behind check 11.
From 1,205 pages scanned since January 2026, share of pages failing each check:
Median Citation Score across the brands measured: 16 out of 100. Roughly half scored zero on citations, never named as a source by any assistant, on any of their own category's buying questions.
Almost none of it tracked with traditional SEO health. Sites with clean architecture, good Core Web Vitals and real backlinks failed extraction at the same rate.
Gemini and Perplexity cite generously. Claude is the strictest gatekeeper of the group. That is not a soft impression. Qwairy analyzed 118,000 AI responses between January and March 2026 and found only 11% of cited domains appeared on more than one platform.
The volume gap is just as wide. Perplexity averages 21.87 citations per response, Google AI Mode 8.34, ChatGPT 7.92. So optimizing for one assistant is not a shortcut. A score from one engine is a rumor.
Two studies found this independently. CXL analyzed 100 Google AI Overview citations. Kevin Indig analyzed 18,012 verified ChatGPT citations pulled from 1.2 million responses. Both landed on position: 44.2% of citations come from the first 30% of a page. Content buried deep in a long post is roughly 2.5x less likely to be cited. Indig calls it the ski ramp effect, a steep cliff after the first third and then a long slow tail.
Methodology: 1,205 pages scanned by the GetJuniors scanner since January 2026. Citation Score is a reproducible 0–100 metric: 10 buyer-intent queries per category, run monthly across ChatGPT, Perplexity, Claude and Gemini; 10 points if cited as a source, 5 if mentioned, 0 if absent. Full method and the complete data set: getjuniors.pro/citation-score. Author: Natasha Brovkina. Last updated 1 August 2026.
The pattern underneath all of it: agents don't reward effort. They reward what's easy to extract, verify and trust.
Undated content reads as less trustworthy. Anonymous pages make weaker citations. If the answer is buried in prose instead of sitting in a direct, extractable block, it is harder to use, so it does not get used.
And llms.txt? Real file, proposed in 2024 by Jeremy Howard, now adopted by major docs platforms including Anthropic's own. But it is a floor, not a strategy. Having it does not make weak content trustworthy.
AI does not reward more content. It rewards trusted knowledge, and trust gets built in a sequence rather than handed out for volume: create exclusive knowledge, turn it into a memorable idea, publish it everywhere consistently, earn trust from real people, and only then does AI cite you.
Most brands try to skip straight to the last step. That is the actual gap.
If you fall in the second group, the checklist above is still free and you lose nothing by reading it.
No. Classic SEO is the foundation, and bad SEO means no AI visibility, but it isn't sufficient. Plenty of technically clean sites score zero citations. They are different readers.
Yes, and it happens constantly. Pull the citations from any AI answer and check them in Ahrefs. Many barely rank. Entering the model's source pool and ranking in search are separate problems.
Your site can be fetched by an AI crawler but produce no usable text, or your business is never named when customers ask your category's buying questions. In our scan, 35.7% of pages were non-extractable and nearly half of brands scored zero citations.
No. It was present on 48.1% of the sites we scanned, including many with zero citations. It's a floor, not a strategy.
Because their source pools barely overlap. Qwairy's analysis of 118,000 responses found only 11% of cited domains appeared on more than one platform. Checking one assistant tells you about roughly one eleventh of your visibility.
Yes. Assistants favour genuine round-ups. A page that only promotes you reads as promotional and gets discounted.
The free scan reads your live site and returns your Citation Score in 30 seconds, then a human-reviewed report by email with every fix above prioritised by severity. It is the same scanner that produced all 1,205 pages of data on this page. No account, no call, no card.
Scan my site free →Sources: GetJuniors scan, 1,205 pages, since January 2026 · Kevin Indig, 18,012 verified ChatGPT citations from 1.2M responses · Qwairy, 118,000 AI responses, Jan–Mar 2026 · Otterly, AI Citation Economy 2026, 1M+ citations · CXL, 100 Google AI Overview citations study · llms.txt, proposed by Jeremy Howard, 2024