AI robots.txt Generator
AI companies send three kinds of bots: crawlers that train models, crawlers that build a search index, and fetchers that open a page when a user asks. They answer to separate robots.txt names, so you can stay visible in AI answers while keeping your content out of training. Pick a policy and copy the file.
Your robots.txt
# AI crawler policy # Generated with https://www.alfaumi.pro/robots-txt-generator User-agent: * Allow: / # AI search crawlers: index pages so answers can link to them (allowed) User-agent: OAI-SearchBot User-agent: Claude-SearchBot User-agent: PerplexityBot User-agent: Meta-WebIndexer Allow: / # User-triggered fetchers: read a page when a person asks about it (allowed) User-agent: ChatGPT-User User-agent: Claude-User User-agent: Perplexity-User User-agent: Meta-ExternalFetcher Allow: / # Training crawlers: collect material for model training (blocked) User-agent: GPTBot User-agent: ClaudeBot User-agent: anthropic-ai User-agent: Google-Extended User-agent: Applebot-Extended User-agent: cohere-training-data-crawler User-agent: CCBot Disallow: /
Which bot does what
The generator covers only bots whose operator documents their purpose. The full table of 54 crawlers, with sources, is on the AI crawlers list.
AI search crawlers
They index pages so that an AI answer can cite and link to them. Block one and its product stops showing your site as a source.
OAI-SearchBot(OpenAI)- Surfaces sites in ChatGPT's search features. OpenAI: sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.
Claude-SearchBot(Anthropic)- Navigates the web to improve Claude's search results. Anthropic says blocking it may reduce your visibility and accuracy in user search results.
PerplexityBot(Perplexity)- Surfaces and links sites in Perplexity search results. Perplexity says it is not used to crawl content for AI foundation models.
Meta-WebIndexer(Meta)- Navigates the web to improve Meta AI search result quality. This is Meta's discoverability token: block it and you stay out of Meta AI answers.
User-triggered fetchers
They fetch one page at the moment a person asks an assistant about it. No crawl, no index: a single request on behalf of a user.
ChatGPT-User(OpenAI)- Fetches a page when a ChatGPT or Custom GPT user asks something. User-initiated, so OpenAI says robots.txt rules may not apply. Its operator says robots.txt may not apply to it.
Claude-User(Anthropic)- Fetches a page when a Claude user asks. Anthropic says blocking it may reduce your visibility for user-directed web search.
Perplexity-User(Perplexity)- Fetches a page when a Perplexity user asks something. Perplexity says this fetcher generally ignores robots.txt, because a user initiated the request. Its operator says robots.txt may not apply to it.
Meta-ExternalFetcher(Meta)- Fetches individual links at a user's request. Meta says it may bypass robots.txt, because the fetch is user-initiated. Its operator says robots.txt may not apply to it.
Training crawlers
They collect material that may train future models. Blocking them is a decision about your content, not about your visibility in AI search.
GPTBot(OpenAI)- Crawls content that may train OpenAI's foundation models. Blocking it keeps content out of training, but does NOT remove the site from ChatGPT search: that is OAI-SearchBot.
ClaudeBot(Anthropic)- Collects web content that may contribute to training Anthropic's models. Blocking it excludes future material from training; search and user fetches are Claude-SearchBot and Claude-User.
anthropic-ai(Anthropic)- Legacy token: Anthropic's current crawler documentation no longer lists it.
Google-Extended(Google)- A robots.txt token, not a crawler: controls Gemini training and grounding in Gemini Apps and Vertex AI. Google says it does not affect inclusion in Google Search. Leaving AI Overviews is a separate Search Console control.
Applebot-Extended(Apple)- Does not crawl. A robots.txt token that opts content out of training Apple's foundation models; pages that disallow it can still appear in search results.
cohere-training-data-crawler(Cohere)- Cohere crawler for collecting training data.
CCBot(Common Crawl)- Common Crawl crawler - nonprofit collecting data for AI research. Used by many AI companies.
After you publish it
Save the file as /robots.txt at the root of your domain, merged with any rules you already have. Then check that reality matches the policy: a CDN or firewall can still turn away a bot your robots.txt welcomes. The bot access checker requests your site as each bot and shows where the two disagree.
Questions
- Does blocking GPTBot remove my site from ChatGPT search?
- No. GPTBot collects training material. ChatGPT search results come from OAI-SearchBot, and pages opened for a user come from ChatGPT-User. You can block the first and allow the other two.
- Why are my private paths repeated in every group?
- A crawler that finds a group naming it follows only that group and ignores the rules under "User-agent: *". Without the repetition, naming a bot would reopen the paths you closed for everyone else.
- Is robots.txt enough to keep an AI bot out?
- It is a published request, not an access control. Well-behaved crawlers follow it, and some user-triggered fetchers are documented by their operators as free to ignore it. To refuse a bot for real, block it at the server or CDN.
- My robots.txt allows a bot, but it still cannot read my site. Why?
- A firewall, CDN or bot-protection rule can refuse the request before robots.txt matters. The policy and the edge are two separate things, and the bot access checker tests both.
AlfaUMi also scores a whole site for search and AI visibility.