Original research

AI knows your brand, not what you do

Google AI cites 85% of sites when asked about them by name. Ask what they actually do, and the figure falls to 28%.

We ran SearchGrade's own audit engine against 800 domains drawn from the Tranco ranking, then asked Google AI two questions about a 123-domain paired subsample. Ask about a site by name and it is cited 84.6% of the time. Ask about what it actually does and that falls to 27.6%.

What did we actually find?

The headline is a gap. Asked “what is x.com and what does it offer?”, Google AI cited the site itself in 84.6% of cases. Asked about the same site's own category, with the brand never mentioned, that fell to 27.6%. Both questions ran in the same record, minutes apart, so each domain is its own control.

  1. 84.6% of sites are cited when asked about by name, and only 27.6% when asked about their own category (n=123 paired domains).
  2. Not one domain was cited for its category but not its name. That cell is exactly 0, across all 123 domains.
  3. Ask what the best product in a category is and Google AI cites youtube.com (46.8%) and reddit.com (36.3%) far more often than the vendors themselves.
  4. AI-crawler blocking is a sector story, not a web story: 4.8% of Government & non-profit sites against 61.4% of News & media, a 12.8x spread.
  5. 138 of 733 domains (18.8%) refused a plain HTTP request and were readable only through a headless browser.

Is a site findable if you don't already know its name?

Mostly not. Of 123 paired domains, 70 were cited by name but not for their category. That is 56.9% of the sample that an AI surfaces only once you already know what to ask for. The reverse cell is empty.

Brand query versus category query, 123 paired domains
34Cited for both70By name only0Category only19Neither

Category citation is a strict subset of brand citation, with no exceptions in 123 domains. The empty cell also validates the design: a query that accidentally leaked its brand would show up precisely as a category-only citation, and none did.

Who wins when you ask for a category?

Not the vendors. Aggregated over the generic questions, Google AI's grounded sources are dominated by user-generated content and review media. reddit.com was cited for 45 other domains' categories.

Most-cited sources across the 124 answered category questions
SourceShare of category questions
youtube.com46.8%
reddit.com36.3%
wikipedia.org12.9%
pcmag.com12.1%
google.com5.6%
cnet.com4.8%
medium.com4.8%
forbes.com4%
Ask an AI what the best product in a category is, and it cites YouTube, Reddit, and review sites. The companies that make the product show up about a quarter of the time.

Which sites block AI crawlers?

Sites control AI access through named user-agents in robots.txt, according to Google's own published crawler documentation. Pooled across the sample, 23% block at least one AI crawler. That figure is accurate and tells the reader almost nothing. Broken out by sector it ranges from 4.8% to 61.4%.

Share of sites blocking at least one AI crawler, by sector
News & media61.4% n=101Social & user-generated56.3% n=48Entertainment & gaming29.4% n=68Education & research17% n=47Other15.1% n=53E-commerce & retail13.3% n=60B2B software (non-developer)12.2% n=49Finance9.5% n=21Telecom & hardware6.1% n=33Developer tools & cloud4.9% n=143Government & non-profit4.8% n=21
The same figures as the chart above, with each sector's denominator
SectorBlocks an AI crawlerSites checked
News & media61.4%101
Social & user-generated56.3%48
Entertainment & gaming29.4%68
Education & research17%47
Other15.1%53
E-commerce & retail13.3%60
B2B software (non-developer)12.2%49
Finance9.5%21
Telecom & hardware6.1%33
Developer tools & cloud4.9%143
Government & non-profit4.8%21

The per-bot split is sharper still

Sites are not adopting a policy toward “AI”. They are blocking specific named bots very unevenly, including two from the same vendor at different rates. The last row is not a bot: it is the blanket User-agent: * floor that catches every crawler a site has not named, and the erratum below explains how it came to be there.

AI crawler blocking rate, per user-agent, n=682
User-agentBlocked by
GPTBot18.6% (127)
ClaudeBot18.2% (124)
Applebot-Extended16.1% (110)
anthropic-ai13.9% (95)
PerplexityBot13% (89)
OAI-SearchBot8.7% (59)
User-agent: * (blanket)E15.7% (39)

Parse rules, so this is checkable: groups follow RFC 9309, inline comments are stripped, and only an exact Disallow: / counts as a block. Anyone can open medium.com/robots.txt and verify the method against a real file.

Erratum E1 · AI crawler blocking, per bot · found 2026-08-28, corrected 2026-08-29

This row was collected under the label Googlebot-Extended. That token does not exist: Google publishes Google-Extended as the robots.txt user-agent for its AI opt-out, so no site can write the name we looked for. All 39 blocks counted here came from a blanket "User-agent: *" rule and none named a Google token, so the row is relabelled to what it in fact measures: the share of sites whose wildcard group carries "Disallow: /", which is the floor under every other row. The genuine Google-Extended opt-out rate is not measured by this study and is not reported here. The live audit was corrected on 2026-08-29. The archived dataset keeps the original label, because it records what was collected.

Does any of this change with a site's rank?

Blocking is concentrated at the head: 26.3% of the top band against 16.3% of the tail, though the decline is not steady and ticks back up in the last band. llms.txt adoption is essentially flat, which is itself a finding: the file is not spreading down the long tail. Citation barely moves at all, spanning 3.5 points across the four bands.

Three headline rates across four Tranco rank bands
0%25%50%75%100%top 1,0381k to 10k10k to 100k100k to 1M
  • Blocks an AI crawler
  • No llms.txt
  • Cited by Google AI
Per-band detail, with the denominator for every rate
BandAnalyzedMedian scoreNo llms.txtBlocks AI botCited by Google AI
top 1,0384795883.5% n=47926.3% n=44880.8% n=448
1k to 10k855883.5% n=8519% n=7982.4% n=85
10k to 100k835689.2% n=8314.7% n=7584.3% n=83
100k to 1M865688.4% n=8616.3% n=8082.6% n=86

How reliable is a single AI query?

We re-audited 30 domains 1.9 hours later. The deterministic checks reproduced perfectly: llms.txt 30/30 identical and the blocked-bot set 30/30 identical. The citation verdict did not, and it moved in one direction only.

Agreement on re-query, split by the original verdict
Was cited, still cited17/17Was not cited, still not9/13

Across the whole sample, Google AI cited 573 of 702 domains (81.6%) when we asked by name. We exclude a further 31, because the model answered those without running a search at all. Those are overwhelmingly household names, the domains likeliest to have been cited had a search run, so dropping them pushes the published rate down. In the paired comparison, 150 records yielded 124 measurable and 123 paired domains; the limitations below account for the rest.

Every flip ran in the same direction: uncited domains became cited, and none went the other way. So 81.6% is a floor, not a point estimate, and a “not cited” result is materially weaker evidence than a “cited” one. We do not describe the remainder as sites that are absent from AI search, because our own data rules that out.

Measure twice and the deterministic checks agree perfectly. The model does not. That gap is the reason we publish this number as a floor.

How did we measure it?

Four steps, in the order they ran. Anyone can repeat them.

  1. Draw the sample. Take Tranco list 5648N and stratify it into four rank bands, drawing the three tail bands with a seeded Fisher-Yates shuffle so those bands are regenerable from the seed alone.
  2. Audit each homepage exactly once. Run all 114 audit modules against the homepage and the domain-root files. Page audits, never a site crawl, because a crawl skips the three domain-level GEO checks this study is about.
  3. Ask Google AI twice. Query Gemini with Google Search grounding once by brand name and once about the site's own category, in the same record minutes apart, so each domain is its own control.
  4. Analyze separately from collection. Write every raw record to disk, then run the analyzers over that directory. Auditing is expensive and non-repeatable; analysis is free and can be re-run as new questions come up.

Terms used above

Tranco rank
A research-oriented domain ranking averaged over 30 days. It ranks DNS and resolver prominence, not visits. This is therefore not a list of “the most visited websites”.
Grounded citation
The domain appears in the source list Google AI returns with its answer. Being named in the answer text is a different thing and is reported separately, never scored.
llms.txt
A proposed plain-text file at a domain's root that gives AI assistants a curated map of the site. We require text/plain, so a single-page-app catch-all route returning HTML does not count as publishing one.
Band
One of four Tranco rank strata: top 1,038, 1k to 10k, 10k to 100k, 100k to 1M. Band 1 is the top 1,038 rather than the top 1,000, because the frame had to be walked past 1,000 to reach 500 qualifying websites.

Frequently asked questions

Does this measure ChatGPT?
No, and we never claim it does. Every citation figure comes from Google's Gemini model (gemini-3.5-flash) with Google Search grounding, run on 2026-08-03. Nothing here queries ChatGPT, and a study that said otherwise would be measuring something it never ran.
What counts as being cited?
The domain appears in the grounded source list Google AI returns alongside its answer. Being named in the answer text does not count. We report that separately and never score it. The distinction matters: Google AI named 128 domains in the text without citing them as a source.
Are the headline percentages precise figures?
No. Treat them as floors. Re-querying 30 domains 1.9 hours later, every cited domain stayed cited, but 4 of 13 uncited domains flipped to cited and no flip ran the other direction. A single grounded query under-detects citation, so 84.6% and 27.6% are both lower bounds. The brand arm sits nearer its ceiling than the category arm, so the gap between them is, if anything, understated.
How did we choose the 800 domains?
We took Tranco list 5648N, generated 2026-07-31, and stratified it into four rank bands. A seeded Fisher-Yates shuffle (seed 20260801) drew the three tail bands, so those 300 domains regenerate exactly from the list ID and the seed. The 500 band-1 domains do not: they came from a published qualification rule applied to 1,038 frame rows, and are reproduced from the label file rather than from the seed.
Which numbers did we decide not to publish?
12 Content and AEO checks have a defect we traced to our own code: they mis-score non-Latin scripts. We therefore withhold every pooled percentage for them. 2 more show a gap we could not attribute to our code, and we do not claim those are bugs. We also withhold grade letters, any sector rate under n=20, and the blocking-to-citation association, which confounding makes unreliable.

Where can I get the data?

Everything behind these numbers is published. The raw per-domain records carry each check's verbatim message, which is where the answers to questions we did not anticipate live.

Sources and further reading

Check your own site

Every check behind this study runs on any URL, free and without an account. It takes about a minute.

Run a free audit on your own site

Questions about the method, or want the raw data? Email us about the study.

Share this study: post it on X, share on LinkedIn, or share on Facebook.