Free GEO Audit

Why AI doesn't recommend your business: a 12-point checklist you can run yourself

A customer types your category into ChatGPT. Best bookkeeping firm for e-commerce. Best plumber in Austin. Whatever it is you sell.

It comes back with three names and one line on each.

You are not one of them.

You did not lose that deal. You were never in the room, and nothing in your analytics will ever tell you it happened. There is no impression, no bounce, no page two. Just an answer that went to someone else.

Short answer: in a scan of 1,205 live pages, 35.7% returned no usable text to an AI agent at all, and almost none of it correlated with SEO health. The checklist below is the one we run. It takes about twenty minutes and needs nothing but a browser.

What the agent actually sees

Two pages side by side. Page A, with content injected client-side, yields 8 words to an agent. Page B, server-rendered, yields 488 words.
Left, what a person sees in a browser. Right, the text actually extracted from the HTML the server returns, which is what a non-executing crawler receives. Same design system, same copy, same page weight. Measured with ai_readability_check.py, extraction is verbatim, not illustrative.

Both pages look finished. Only one of them exists to an AI agent.

The difference is not quality, length or design. It is where the content gets assembled. Page A builds itself in the browser, so a crawler that does not execute JavaScript receives navigation and nothing else. Page B ships the same words in the HTML.

Eight words against 488. That is the whole problem in one image, and it is the failure nothing in your analytics reports.

The checklist

Block A · fix these first

Can a machine read the page at all?

Check 01

Text is in the HTML, not painted into images

Open the page, hit Ctrl+U, and search for a sentence you can see on screen. If it isn't in the source, an agent can't read it either.

Check 02

The main answer survives with JavaScript off

Disable JS and reload. If the page goes blank, many crawlers see the same blank page.

Check 03

The page isn't blocked

Check robots.txt for the AI crawlers specifically: GPTBot, ClaudeBot, PerplexityBot, Google-Extended. Blocking them is usually accidental, inherited from a template.

Check 04

The page has more than boilerplate

If the unique text is under roughly 150 words, there is nothing worth extracting. Half a service page of navigation and a contact form counts as empty.

Block B · cheap, and skipped by four sites in five

Can it tell who is speaking, and when?

Check 05

A visible publication or update date

77.1% of the pages we scanned had no date anywhere on the content. Undated reads as untrustworthy, especially in fast-moving categories.

Check 06

A named author with a real bio

Only 21.2% of pages had a visible author. "Admin" and "The Team" do not count. An assistant needs to know who is making a claim before it repeats it.

Check 07

Organization schema that resolves to one entity

Same legal name, same address, same handles everywhere. If your name renders three different ways across the web, an agent can't merge you into a single thing it can recommend.

Block C · shape

Is the answer in a shape a machine can lift?

Check 08

The first two sentences answer who, what and for whom

This is the ten-second test. Open your own service page and read them. Do they answer the buyer's question, or describe your journey since 1998? An assistant cannot quote the second one.

Check 09

A direct answer block near the top

44.2% of AI citations come from the first 30% of a page (Kevin Indig, 18,012 verified ChatGPT citations). Put the extractable answer there, not in paragraph nine.

Check 10

A real FAQ section with FAQ schema

62.8% of pages we scanned had none. FAQ blocks are the one exception to front-loading, because each question and answer is self-contained, so they get cited even at the bottom of a long page.

Block D · placement

Are you where the model already looks?

Check 11

You appear in third-party sources, not just on your own site

Reviews, directories, community discussion, round-ups. Assistants recommend businesses they see named by someone other than the business.

Check 12

You are placed where your assistants actually read

Each model has a different source diet, and the placement table further down maps it. A decent page on a platform the model already reads beats an excellent page where it never looks.

Now check whether it worked

This is the loop we run monthly. It takes ten minutes per query and it is the only honest way to know.

Step 01

Pick a real buyer query

Not your brand name, the question a customer types before hiring. "Best plumber in [city]", "best local rank tracker", "who should I hire for X".

Step 02

Ask it in ChatGPT, Claude and Perplexity

Separately. They do not agree with each other.

Step 03

Check whether you are named at all

Mentioned in the text is one thing. Cited as a source with a link is another, and worth more.

Step 04

Open the citations

Pull the source URLs from the footnotes or sidebar.

Step 05

Run those URLs through Ahrefs or SimilarWeb

This is the step that surprises people. Many cited sources barely rank in Google. Ranking well and getting cited are not the same channel.

Step 06

If you are absent

Add or update a listicle in your category, build mentions on that model's preferred platforms, and sharpen structure and markup on the page you want cited.

Repeat monthly against the same queries. Changing the queries each time makes the result unmeasurable.

Where each assistant looks

Verified by hand across our own tracked queries. This is observation, not a controlled study, so treat it as a starting map rather than statistics.

AssistantLeans onSo publish
ChatGPTWikipedia, news agencies, government sites, long reviews and listiclesown-site listicles under buyer-intent queries; encyclopedic and editorial placements
Claudereviews and local sites, plus a broad mix of blogs and guidesdirectories and review platforms; detailed service and neighbourhood pages
Perplexity, GeminiYouTube, Reddit, large editorial sitesYouTube with real descriptions and timestamps; genuine Reddit participation

One counter-intuitive rule for listicles: include your competitors. Assistants reward genuine round-ups and discount self-promotion. A comparison page that names rivals outperforms a page that only sells you.

How this squares with the published numbers. Otterly's 2026 analysis of more than a million citations reports a different cut of the same territory: news and media account for 20 to 30% of citations depending on platform, community forums for 5.9 to 16.9%, and brand sites get mentioned often but linked rarely.

Those figures do not contradict the table above. They answer a different question. Otterly measures what share of all citations each source type earns across a broad query set. The table records which source types actually turned up against a specific set of buyer-intent queries in service categories. A source type can be a small slice of the internet-wide total and still be the one that decides your category.

Where the two agree is the part that matters here: both put community and review platforms well ahead of brand-owned pages as citation sources. That is the whole argument behind check 11.

The evidence behind the checklist

From 1,205 pages scanned since January 2026, share of pages failing each check:

No visible author78.8%
No date on the content77.1%
No FAQ schema62.8%
No llms.txt (measured per site, not per page)51.9%
No direct answer block50.2%
Thin or non-extractable35.7%

Median Citation Score across the brands measured: 16 out of 100. Roughly half scored zero on citations, never named as a source by any assistant, on any of their own category's buying questions.

Almost none of it tracked with traditional SEO health. Sites with clean architecture, good Core Web Vitals and real backlinks failed extraction at the same rate.

The engines don't agree with each other either

Gemini and Perplexity cite generously. Claude is the strictest gatekeeper of the group. That is not a soft impression. Qwairy analyzed 118,000 AI responses between January and March 2026 and found only 11% of cited domains appeared on more than one platform.

The volume gap is just as wide. Perplexity averages 21.87 citations per response, Google AI Mode 8.34, ChatGPT 7.92. So optimizing for one assistant is not a shortcut. A score from one engine is a rumor.

Position inside the page decides a lot

Two studies found this independently. CXL analyzed 100 Google AI Overview citations. Kevin Indig analyzed 18,012 verified ChatGPT citations pulled from 1.2 million responses. Both landed on position: 44.2% of citations come from the first 30% of a page. Content buried deep in a long post is roughly 2.5x less likely to be cited. Indig calls it the ski ramp effect, a steep cliff after the first third and then a long slow tail.

Methodology: 1,205 pages scanned by the GetJuniors scanner since January 2026. Citation Score is a reproducible 0–100 metric: 10 buyer-intent queries per category, run monthly across ChatGPT, Perplexity, Claude and Gemini; 10 points if cited as a source, 5 if mentioned, 0 if absent. Full method and the complete data set: getjuniors.pro/citation-score. Author: Natasha Brovkina. Last updated 1 August 2026.

Why it's broken

The pattern underneath all of it: agents don't reward effort. They reward what's easy to extract, verify and trust.

Undated content reads as less trustworthy. Anonymous pages make weaker citations. If the answer is buried in prose instead of sitting in a direct, extractable block, it is harder to use, so it does not get used.

And llms.txt? Real file, proposed in 2024 by Jeremy Howard, now adopted by major docs platforms including Anthropic's own. But it is a floor, not a strategy. Having it does not make weak content trustworthy.

AI does not reward more content. It rewards trusted knowledge, and trust gets built in a sequence rather than handed out for volume: create exclusive knowledge, turn it into a memorable idea, publish it everywhere consistently, earn trust from real people, and only then does AI cite you.

Most brands try to skip straight to the last step. That is the actual gap.

Is this worth your time?

Worth running if

  • You sell something people research before they buy, and a shortlist decides it
  • You already rank reasonably in Google and still hear "we found someone else"
  • You are willing to change pages you already considered finished, and to wait sixty days before judging

Not worth running if

  • You want a plugin that fixes this in an afternoon. There isn't one, and llms.txt is not it. 48.1% of the sites we scanned already had one
  • Your pipeline is paid and referral, and you're happy with that
  • You expect this to lift your Google rankings. Usually it doesn't, and that separation is the entire finding

If you fall in the second group, the checklist above is still free and you lose nothing by reading it.

FAQ

Does good SEO make me visible in ChatGPT?

No. Classic SEO is the foundation, and bad SEO means no AI visibility, but it isn't sufficient. Plenty of technically clean sites score zero citations. They are different readers.

Can a page be cited by AI without ranking in Google?

Yes, and it happens constantly. Pull the citations from any AI answer and check them in Ahrefs. Many barely rank. Entering the model's source pool and ranking in search are separate problems.

What does it mean to be AI-invisible?

Your site can be fetched by an AI crawler but produce no usable text, or your business is never named when customers ask your category's buying questions. In our scan, 35.7% of pages were non-extractable and nearly half of brands scored zero citations.

Is adding llms.txt enough?

No. It was present on 48.1% of the sites we scanned, including many with zero citations. It's a floor, not a strategy.

Why do different assistants recommend different businesses?

Because their source pools barely overlap. Qwairy's analysis of 118,000 responses found only 11% of cited domains appeared on more than one platform. Checking one assistant tells you about roughly one eleventh of your visibility.

Should I list competitors on my own comparison page?

Yes. Assistants favour genuine round-ups. A page that only promotes you reads as promotional and gets discounted.

Run it against your own site

The free scan reads your live site and returns your Citation Score in 30 seconds, then a human-reviewed report by email with every fix above prioritised by severity. It is the same scanner that produced all 1,205 pages of data on this page. No account, no call, no card.

Scan my site free →

Sources: GetJuniors scan, 1,205 pages, since January 2026 · Kevin Indig, 18,012 verified ChatGPT citations from 1.2M responses · Qwairy, 118,000 AI responses, Jan–Mar 2026 · Otterly, AI Citation Economy 2026, 1M+ citations · CXL, 100 Google AI Overview citations study · llms.txt, proposed by Jeremy Howard, 2024