Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers
The AI Visibility Index separates "being visible to AI" into two distinct issues: answer presence and machine readability. Between July 23 and August 2, 2026, 193 sites were used in 5,978 AI assistant queries, and 94.8% of them were never named by ChatGPT, Claude, Gemini, or Perplexity. Only 1.9% of all answers mentioned the business asked about. Meanwhile, just 8.9% of 436 sites block at least one AI crawler in robots.txt, with blocking aimed mainly at training crawlers like GPTBot rather than the retrieval agents used for live answers.
Our Terms and Privacy Policy were updated on July 27, 2026. Read the Terms and the Privacy Policy.
94.8% of audited sites are never named in an AI answer
1.9% of 5,978 assistant answers named the business asked about
8.9% block at least one AI crawler in robots.txt
What the AI Visibility Index measures
The Index tracks two things separately, because "being visible to AI" is really two problems with two different causes.
Answer presence
A customer asks an assistant for the best provider of a category in a place, and the assistant answers with a handful of names. This pillar asks whether the audited business is one of those names. Every audited business is queried against ChatGPT, Claude, Gemini and Perplexity, and each answer is matched against the audited business.
Machine readability
This pillar is what an assistant's crawler finds on arrival: whether robots.txt lets it in, whether the site publishes a sitemap, whether the homepage states the business's identity and location in machine-readable form. Every one of these signals is first-party, read by our own crawler directly from the site.
Presence is the outcome; readability is the input under a site owner's control. A site can be perfectly marked up and still go unnamed, which is why both are published side by side rather than combined into a single score. The signals come from the same free scan anyone can run on their own site.
Key findings
94.8%
of audited sites never surface in an AI answer
Only 10 of 193 sites were named even once across 5,978 assistant answers to buying-intent questions.
2.9%
Gemini has the widest reach in the panel
Gemini named the audited business in 2.9% of its answers, ahead of every other assistant in the panel. All four were asked the same questions about the same businesses on the same days, so the gap is a difference between assistants rather than a difference between questions.
8.9%
block at least one AI crawler in robots.txt
39 of 436 sites disallow an AI user-agent outright, most of them without distinguishing between crawlers that train models and agents that fetch a page to answer a live question.
38
sites block GPTBot; only 4 block OAI-SearchBot
Blocking is aimed at training crawlers rather than at retrieval agents. GPTBot is disallowed by 38 sites and ClaudeBot by 36, while the agents that actually fetch a page to answer a live question, OAI-SearchBot and Perplexity-User, are disallowed by 4 sites and 3 sites. A site that blocks the whole group pays for it in lost citations.
54.6%
publish any schema.org structured data
238 of 436 sites carry structured data on their homepage. The rest leave an assistant nothing to read but prose.
19.3%
publish LocalBusiness schema
The markup that states a business's name, address and category in machine-readable form is the least-adopted signal measured here, present on 84 sites. Most of this corpus is local businesses.
The data
Share of answers naming the audited business
2026-07-23 to 2026-08-02 • 193 sites • 5,978 answers
Gemini 2.9% 8 of 193 sites
ChatGPT 1.7% 5 of 193 sites
Claude 1.6% 5 of 193 sites
Perplexity 1.6% 8 of 193 sites
Every assistant was asked the same questions about the same businesses on the same days, so the differences below are differences between assistants. ChatGPT answered 1,352 of its 1,544 queries; the 192 that failed at the provider are excluded from its denominator rather than counted as absences.
AI crawlers blocked in robots.txt
Share of 436 audited sites disallowing each user-agent
CCBot 8.7% 38 sites - training
GPTBot 8.7% 38 sites - training
Bytespider 8.5% 37 sites - training
ClaudeBot 8.3% 36 sites - training
Google-Extended 8.3% 36 sites - training
anthropic-ai 1.4% 6 sites - training
PerplexityBot 1.1% 5 sites - retrieval
ChatGPT-User 0.9% 4 sites - retrieval
OAI-SearchBot 0.9% 4 sites - retrieval
Perplexity-User 0.7% 3 sites - retrieval
Training crawlers harvest pages to build models; retrieval agents fetch a page in response to a live question. Blocking the first costs a site nothing in today's answers. Blocking the second makes the site uncitable, a far larger consequence than the gap between these two columns.
AI-readiness signals found on audited sites
2026-05-16 to 2026-08-02 • 436 sites • latest audit per domain
Publishes robots.txt 79.4% 346 of 436
Publishes sitemap.xml 73.9% 322 of 436
Points robots.txt at sitemap 67.9% 296 of 436
Has a meta description 77.1% 336 of 436
Has Open Graph tags 69.5% 303 of 436
Has any schema.org data 54.6% 238 of 436
Has LocalBusiness schema 19.3% 84 of 436
Blocks at least one AI crawler 8.9% 39 of 436
Methodology
The dataset
Every figure comes from audits run through Website Auditor's free scanner between 2026-03-30 and 2026-08-02: 531 stored audits across 458 distinct domains, in 38 detected sectors. Aggregates are computed over Website Auditor's reports, the analytics view that excludes our own parent company auditing its own properties, so internal traffic cannot inflate a share. The aggregate itself is committed as a SQL function and re-run unchanged for every edition, so one edition is a comparison against the last rather than a fresh set of judgement calls.
How each number is calculated
One domain, one vote. Both pillars use the most recent audit per domain. Counting audits instead would give a site scanned nine times nine times the influence over every share.
Answer presence covers 2026-07-23 to 2026-08-02: 193 sites and 5,978 assistant answers. A site counts as present if it was named in at least one answer from at least one assistant.
Machine readability covers 2026-05-16 to 2026-08-02 across 436 sites, a longer window and a larger sample, because these signals are read by our own crawler and depend on no third-party provider.
Failed queries are excluded, not counted as absences. A provider timeout is a fact about the provider, not evidence about the business. Per-assistant failure counts are published above, so a smaller denominator is something a reader can see rather than an inference.
Every percentage is a derivation. The published dataset holds integer counts only, and every share on this page is computed from two of them at render time. A headline therefore cannot drift away from its own data.
Assumptions and limitations
Every figure above comes with conditions, and these are the conditions. They are published with the same weight as the findings.
The corpus is self-selected, not a random sample of the web: every site in it is one whose owner chose to run a free audit, which skews toward small and mid-sized businesses actively working on their web presence. Treat the shares as descriptive of that population, not of all websites.
Cross-assistant figures start on 2026-07-23 because earlier audits fanned a single Perplexity answer across the ChatGPT, Claude and Gemini cards. Those earlier rows are real answers, but all of them are Perplexity's, so comparing assistants over them would measure one provider against its own output.
ChatGPT answered 1,352 of its 1,544 panel queries; the 192 that failed at the provider are excluded from its denominator rather than counted as absences. Its rate therefore rests on a smaller sample than every other assistant in the panel.
Appearance is measured by matching the audited business against the businesses named in an assistant's answer. Name matching is fuzzy, and a business identified only by its domain, roughly one in five, is the hardest case. Presence is therefore more likely to be an under-count than an over-count.
Assistants are non-deterministic, and their indexes are in constant motion. These figures describe the answers given inside the stated window; the same queries run after 2026-08-02 will produce different answers.
Crawl signals are read from three places only: the homepage, robots.txt and sitemap.xml. A site carrying structured data on its inner pages but not on its homepage is counted here as a site with no structured data.
Citation and reuse
Figures are free to reproduce with attribution and a link to this page. For the underlying methodology, or a cut of the data for a specific sector, write to [email protected] with your question.
Update history
This page is the permanent home of the Index. Each new edition updates the figures above and adds a dated row here rather than moving to a new URL, so links and citations keep resolving to current data.
2026-08-02
Second edition. The panel grew from 175 sites to 193 sites and now runs to 2026-08-02, and the crawl-signal corpus grew from 418 sites to 436 sites. Every headline share moved by less than a point: a slightly smaller share of sites named in an answer, and a slightly smaller share blocking an AI crawler. Gemini keeps the panel's widest reach.
2026-07-31
First edition. Cross-assistant presence measured over 2026-07-23 to 2026-07-31, the first period in which ChatGPT, Claude, Gemini and Perplexity were each queried through their own provider. Crawl signals measured across 418 sites audited since 2026-05-16, at one audit per domain.
Where does your site sit in this data?
The Index is built from the same free audit anyone can run on their own site. It checks the AI-readiness signals measured above and queries all four assistants for your business by name, with no signup and a result in about a minute.
Run a free AI visibility audit