Skip to content
website.show

AI crawler checker

Enter a site and see which AI bots its robots.txt lets read the homepage: the ones that fetch pages for answers, and the ones that collect them for training. We read one file, the site’s robots.txt.

How the file is read

The way RFC 9309, the robots.txt standard, says crawlers read it: a bot follows the group that names it and falls back to * only when none does; when an Allow and a Disallow both match, the longer rule wins, and Allow wins a tie. A missing file means every crawler may read the site; a server error means crawlers treat the whole site as closed.

What a blocked bot costs

Blocking a search bot, OAI-SearchBot, Claude-SearchBot or PerplexityBot, can keep your pages out of that assistant’s answers. Blocking a training bot, GPTBot, ClaudeBot or Google-Extended, keeps new pages out of future models and leaves the answers alone. Google says Google-Extended does not affect Google Search.

The file is not the whole story: a CDN or firewall can turn bots away before they reach it, and some fetches a person triggers do not follow it. Our guide, Is your site blocking ChatGPT?, covers each case with the vendors’ documentation.