When someone asks ChatGPT or Perplexity for a tool like yours, the answer is built from pages a bot fetched. If your robots.txt turns that bot away, your product is not one of the pages the answer can quote, however good the page is.
Most sites that block AI crawlers never decided to. The rule came with a template, a security plugin or a CDN setting. This guide covers which bots matter, what blocking each one costs, and how to check your own file. It is based on the crawler documentation of OpenAI, Anthropic, Perplexity and Google, and on the robots.txt standard, as of October 2026.
Which bots fetch your pages for AI answers?
Each company runs separate bots for separate jobs: one fetches pages for its search and answers, one collects pages for training future models, and one visits a page when a user asks about it. robots.txt can treat each of them differently.

| Company | Answers and search | Model training | On a user's request |
|---|---|---|---|
| OpenAI | OAI-SearchBot | GPTBot | ChatGPT-User |
| Anthropic | Claude-SearchBot | ClaudeBot | Claude-User |
| Perplexity | PerplexityBot | Not used for training | Perplexity-User |
Googlebot | Google-Extended |
Google is the odd one out. Google-Extended is a token you can name in robots.txt, not a separate crawler, and Google says it does not affect a site's inclusion in Google Search or its ranking there. Pages reach Google Search, AI Overviews included, through Googlebot.
What happens when you block each one?
Blocking a search bot can take you out of that assistant's answers. Blocking a training bot keeps your new pages out of future models and leaves the answers alone.
- Blocking a search bot can cost you answers. OpenAI says sites that opt out of
OAI-SearchBotwill not be shown in ChatGPT search answers, though they can still appear as navigational links. Perplexity usesPerplexityBotto surface and link websites in its results, and Anthropic usesClaude-SearchBotto index content for search. - Blocking a training bot affects future models, not answers. Disallowing
GPTBottells OpenAI the content should not be used to train its models. BlockingClaudeBotexcludes your future pages from Anthropic's training data.Google-Extendedcovers Gemini training and grounding in Gemini apps and Vertex AI. - A fetch that a user triggers visits a page because a person asked about it. Anthropic says
Claude-Userrespects robots.txt. OpenAI says robots.txt rules may not apply toChatGPT-User, and Perplexity saysPerplexity-Usergenerally ignores them.
So keeping AI training out costs nothing in search answers, as long as the search bots can still get in. Blocking Googlebot is a different decision: it takes you out of Google Search altogether.
How do you check your robots.txt in two minutes?
Open yoursite.com/robots.txt in a browser and read it top to bottom.
- Look for a
User-agent: *group withDisallow: /. That one line closes the site to every crawler that has no group of its own, AI search bots included. - Look for each bot by name:
OAI-SearchBot,Claude-SearchBot,PerplexityBot,GPTBot,ClaudeBot,Google-Extended. ADisallow: /under a name blocks that bot from the whole site. - Check what your CDN and security tools add. Some serve a robots.txt of their own in front of yours, or block bots before they reach the file.
Every product listed on website.show gets this check on its listing page. It reads the site's robots.txt the way the standard says crawlers should, and reports which of those six bots are blocked. The same check runs weekly, and our crawler page lists what we fetch and when.
Why can a rule for one bot open pages you meant to close?
Because a crawler follows only the group that names it. The robots.txt standard, RFC 9309, says a crawler obeys the group whose user-agent matches it and falls back to the * group only when no group matches.

This catches people out when they add a line to welcome one AI bot. If the * group closes /admin/ and /api/, and you add User-agent: OAI-SearchBot with Allow: /, that bot no longer reads the * rules, so those paths are open to it. Repeat the shared Disallow lines inside every named group. Our own robots.txt does exactly that for the AI bots it names.
Two smaller rules from the same standard help when you read a file. When an Allow and a Disallow both match a path, the longer, more specific rule wins. When they match equally, Allow wins.
Which robots.txt fits a product site?
For most product sites, one of two files. The first lets every crawler in and closes only what nobody should index:
User-agent: *
Disallow: /api/
Disallow: /admin/
Sitemap: https://yoursite.com/sitemap.xmlThe second keeps AI training out and leaves search and answers open. The training bots get a group of their own:
User-agent: *
Disallow: /api/
Disallow: /admin/
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
Sitemap: https://yoursite.com/sitemap.xmlIn the second file the training bots are closed out entirely, so the shared lines add nothing for them. Search bots have no group of their own and follow *.
Can your CDN block AI bots without telling you?
It can, and it is worth checking even if you never changed a setting. Cloudflare offers a managed robots.txt that, when you turn it on, puts Disallow rules for GPTBot, ClaudeBot, Google-Extended, CCBot and several other AI bots in front of your own file. It does not list OAI-SearchBot, Claude-SearchBot or PerplexityBot, so turning it on does not take you out of ChatGPT, Claude or Perplexity answers.
Cloudflare can also block AI bots outright, whatever robots.txt says. Its documentation describes new defaults for new domains from September 15, 2026: on pages that display ads, bots classed as training or agent bots are blocked, and so are crawlers that mix search with training, while search bots stay allowed. The older "Block AI bots" switch is being retired in favor of these separate policies. If your site is behind Cloudflare, check its AI crawler settings as well as the file.
What else stops an AI crawler from reading a page?
JavaScript can, because most AI crawlers read only the HTML your server sends. Vercel measured the major AI crawlers in December 2024 and found that none of them render JavaScript, apart from Gemini, which uses Googlebot, and AppleBot. A page that builds its text in the browser can look empty to them even when robots.txt lets them in.
Open your page with JavaScript turned off, or view its source, and look for your headline and the first paragraphs. If they are not in the HTML your server sends, assistants may have nothing to quote. The site check on website.show listings tests this too.
FAQ
Does blocking GPTBot take my site out of ChatGPT?
No, because GPTBot only collects pages for training. ChatGPT's search answers come from pages fetched by OAI-SearchBot, so a site that blocks GPTBot and allows OAI-SearchBot can still be shown in those answers.
Does Google-Extended affect AI Overviews?
Not according to Google, which says Google-Extended controls whether content is used for Gemini training and grounding and does not affect a site's inclusion or ranking in Google Search. AI Overviews are part of Google Search, which Googlebot crawls.
Do AI bots always obey robots.txt?
The crawlers in this guide say they do, and Anthropic says all three of its bots honor it. The exceptions are fetches a user triggers: OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them.
Is llms.txt a replacement for robots.txt?
No, llms.txt is a proposed file that gives assistants a short map of your site in Markdown. It does not grant or refuse access; robots.txt still decides which bots may fetch which pages.
Should I let AI companies train on my site?
It is your call, and it is separate from appearing in answers. Keeping training bots out does not stop search bots from fetching your pages, so you can stay quotable in AI answers either way.
Sources
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity: Perplexity crawlers
- Google: Google's common crawlers
- RFC 9309: Robots Exclusion Protocol
- Cloudflare: Managed robots.txt
- Cloudflare: Block AI bots
- Vercel: The rise of the AI crawler
- Where to launch your product in 2026
Every bot name and rule above comes from the source linked, read on October 6, 2026. Crawler documentation changes, so check the vendor's page before you rely on a detail.