AI knows your brand, not what you do
We ran SearchGrade's own audit engine against 800 domains drawn from the Tranco ranking, then asked Google AI two questions about a 123-domain paired subsample. Ask about a site by name and it is cited 84.6% of the time. Ask about what it actually does and that falls to 27.6%.
What did we actually find?
The headline is a gap. Asked “what is x.com and what does it offer?”, Google AI cited the site itself in 84.6% of cases. Asked about the same site's own category, with the brand never mentioned, that fell to 27.6%. Both questions ran in the same record, minutes apart, so each domain is its own control.
- 84.6% of sites are cited when asked about by name, and only 27.6% when asked about their own category (n=123 paired domains).
- Not one domain was cited for its category but not its name. That cell is exactly 0, across all 123 domains.
- Ask what the best product in a category is and Google AI cites youtube.com (46.8%) and reddit.com (36.3%) far more often than the vendors themselves.
- AI-crawler blocking is a sector story, not a web story: 4.8% of Government & non-profit sites against 61.4% of News & media, a 12.8x spread.
- 138 of 733 domains (18.8%) refused a plain HTTP request and were readable only through a headless browser.
Is a site findable if you don't already know its name?
Mostly not. Of 123 paired domains, 70 were cited by name but not for their category. That is 56.9% of the sample that an AI surfaces only once you already know what to ask for. The reverse cell is empty.
Category citation is a strict subset of brand citation, with no exceptions in 123 domains. The empty cell also validates the design: a query that accidentally leaked its brand would show up precisely as a category-only citation, and none did.
Who wins when you ask for a category?
Not the vendors. Aggregated over the generic questions, Google AI's grounded sources are dominated by user-generated content and review media. reddit.com was cited for 45 other domains' categories.
| Source | Share of category questions |
|---|---|
youtube.com | 46.8% |
reddit.com | 36.3% |
wikipedia.org | 12.9% |
pcmag.com | 12.1% |
google.com | 5.6% |
cnet.com | 4.8% |
medium.com | 4.8% |
forbes.com | 4% |
Ask an AI what the best product in a category is, and it cites YouTube, Reddit, and review sites. The companies that make the product show up about a quarter of the time.
Which sites block AI crawlers?
Sites control AI access through named user-agents in robots.txt, according to Google's own published crawler documentation. Pooled across the sample, 23% block at least one AI crawler. That figure is accurate and tells the reader almost nothing. Broken out by sector it ranges from 4.8% to 61.4%.
| Sector | Blocks an AI crawler | Sites checked |
|---|---|---|
| News & media | 61.4% | 101 |
| Social & user-generated | 56.3% | 48 |
| Entertainment & gaming | 29.4% | 68 |
| Education & research | 17% | 47 |
| Other | 15.1% | 53 |
| E-commerce & retail | 13.3% | 60 |
| B2B software (non-developer) | 12.2% | 49 |
| Finance | 9.5% | 21 |
| Telecom & hardware | 6.1% | 33 |
| Developer tools & cloud | 4.9% | 143 |
| Government & non-profit | 4.8% | 21 |
The per-bot split is sharper still
Sites are not adopting a policy toward “AI”. They are blocking specific named bots very unevenly, including two from the same vendor at different rates. The last row is not a bot: it is the blanket User-agent: * floor that catches every crawler a site has not named, and the erratum below explains how it came to be there.
| User-agent | Blocked by |
|---|---|
GPTBot | 18.6% (127) |
ClaudeBot | 18.2% (124) |
Applebot-Extended | 16.1% (110) |
anthropic-ai | 13.9% (95) |
PerplexityBot | 13% (89) |
OAI-SearchBot | 8.7% (59) |
User-agent: * (blanket)E1 | 5.7% (39) |
Parse rules, so this is checkable: groups follow RFC 9309, inline comments are stripped, and only an exact Disallow: / counts as a block. Anyone can open medium.com/robots.txt and verify the method against a real file.
Erratum E1 · AI crawler blocking, per bot · found 2026-08-28, corrected 2026-08-29
This row was collected under the label Googlebot-Extended. That token does not exist: Google publishes Google-Extended as the robots.txt user-agent for its AI opt-out, so no site can write the name we looked for. All 39 blocks counted here came from a blanket "User-agent: *" rule and none named a Google token, so the row is relabelled to what it in fact measures: the share of sites whose wildcard group carries "Disallow: /", which is the floor under every other row. The genuine Google-Extended opt-out rate is not measured by this study and is not reported here. The live audit was corrected on 2026-08-29. The archived dataset keeps the original label, because it records what was collected.
Does any of this change with a site's rank?
Blocking is concentrated at the head: 26.3% of the top band against 16.3% of the tail, though the decline is not steady and ticks back up in the last band. llms.txt adoption is essentially flat, which is itself a finding: the file is not spreading down the long tail. Citation barely moves at all, spanning 3.5 points across the four bands.
- Blocks an AI crawler
- No llms.txt
- Cited by Google AI
| Band | Analyzed | Median score | No llms.txt | Blocks AI bot | Cited by Google AI |
|---|---|---|---|---|---|
| top 1,038 | 479 | 58 | 83.5% n=479 | 26.3% n=448 | 80.8% n=448 |
| 1k to 10k | 85 | 58 | 83.5% n=85 | 19% n=79 | 82.4% n=85 |
| 10k to 100k | 83 | 56 | 89.2% n=83 | 14.7% n=75 | 84.3% n=83 |
| 100k to 1M | 86 | 56 | 88.4% n=86 | 16.3% n=80 | 82.6% n=86 |
How reliable is a single AI query?
We re-audited 30 domains 1.9 hours later. The deterministic checks reproduced perfectly: llms.txt 30/30 identical and the blocked-bot set 30/30 identical. The citation verdict did not, and it moved in one direction only.
Across the whole sample, Google AI cited 573 of 702 domains (81.6%) when we asked by name. We exclude a further 31, because the model answered those without running a search at all. Those are overwhelmingly household names, the domains likeliest to have been cited had a search run, so dropping them pushes the published rate down. In the paired comparison, 150 records yielded 124 measurable and 123 paired domains; the limitations below account for the rest.
Every flip ran in the same direction: uncited domains became cited, and none went the other way. So 81.6% is a floor, not a point estimate, and a “not cited” result is materially weaker evidence than a “cited” one. We do not describe the remainder as sites that are absent from AI search, because our own data rules that out.
Measure twice and the deterministic checks agree perfectly. The model does not. That gap is the reason we publish this number as a floor.
How did we measure it?
Four steps, in the order they ran. Anyone can repeat them.
- Draw the sample. Take Tranco list 5648N and stratify it into four rank bands, drawing the three tail bands with a seeded Fisher-Yates shuffle so those bands are regenerable from the seed alone.
- Audit each homepage exactly once. Run all 114 audit modules against the homepage and the domain-root files. Page audits, never a site crawl, because a crawl skips the three domain-level GEO checks this study is about.
- Ask Google AI twice. Query Gemini with Google Search grounding once by brand name and once about the site's own category, in the same record minutes apart, so each domain is its own control.
- Analyze separately from collection. Write every raw record to disk, then run the analyzers over that directory. Auditing is expensive and non-repeatable; analysis is free and can be re-run as new questions come up.
Terms used above
- Tranco rank
- A research-oriented domain ranking averaged over 30 days. It ranks DNS and resolver prominence, not visits. This is therefore not a list of “the most visited websites”.
- Grounded citation
- The domain appears in the source list Google AI returns with its answer. Being named in the answer text is a different thing and is reported separately, never scored.
- llms.txt
- A proposed plain-text file at a domain's root that gives AI assistants a curated map of the site. We require
text/plain, so a single-page-app catch-all route returning HTML does not count as publishing one. - Band
- One of four Tranco rank strata: top 1,038, 1k to 10k, 10k to 100k, 100k to 1M. Band 1 is the top 1,038 rather than the top 1,000, because the frame had to be walked past 1,000 to reach 500 qualifying websites.
Frequently asked questions
- Does this measure ChatGPT?
- No, and we never claim it does. Every citation figure comes from Google's Gemini model (gemini-3.5-flash) with Google Search grounding, run on 2026-08-03. Nothing here queries ChatGPT, and a study that said otherwise would be measuring something it never ran.
- What counts as being cited?
- The domain appears in the grounded source list Google AI returns alongside its answer. Being named in the answer text does not count. We report that separately and never score it. The distinction matters: Google AI named 128 domains in the text without citing them as a source.
- Are the headline percentages precise figures?
- No. Treat them as floors. Re-querying 30 domains 1.9 hours later, every cited domain stayed cited, but 4 of 13 uncited domains flipped to cited and no flip ran the other direction. A single grounded query under-detects citation, so 84.6% and 27.6% are both lower bounds. The brand arm sits nearer its ceiling than the category arm, so the gap between them is, if anything, understated.
- How did we choose the 800 domains?
- We took Tranco list 5648N, generated 2026-07-31, and stratified it into four rank bands. A seeded Fisher-Yates shuffle (seed 20260801) drew the three tail bands, so those 300 domains regenerate exactly from the list ID and the seed. The 500 band-1 domains do not: they came from a published qualification rule applied to 1,038 frame rows, and are reproduced from the label file rather than from the seed.
- Which numbers did we decide not to publish?
- 12 Content and AEO checks have a defect we traced to our own code: they mis-score non-Latin scripts. We therefore withhold every pooled percentage for them. 2 more show a gap we could not attribute to our code, and we do not claim those are bugs. We also withhold grade letters, any sector rate under n=20, and the blocking-to-citation association, which confounding makes unreliable.
Where can I get the data?
Everything behind these numbers is published. The raw per-domain records carry each check's verbatim message, which is where the answers to questions we did not anticipate live.
- Download the full dataset archive. It holds the raw records, the sample definition, the sector taxonomy, and every category query.
- Citable archive with a DOI. The record at 10.5281/zenodo.21885506 is permanently archived and independent of this site.
Sources and further reading
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation. The exact list used here, 5648N.
- RFC 9309: Robots Exclusion Protocol. The parsing rules our robots.txt reader follows.
- Generative engine optimization. Background on the field this study measures.
- The llms.txt proposal. The specification for the file we tested for.
- SearchGrade scoring methodology. How every check in this study is scored and weighted.
Check your own site
Every check behind this study runs on any URL, free and without an account. It takes about a minute.
Run a free audit on your own siteQuestions about the method, or want the raw data? Email us about the study.
Share this study: post it on X, share on LinkedIn, or share on Facebook.