How to Configure robots.txt for AI Crawlers
AI companies run separate bots for search, training and user requests. Here's how to tell them apart, three robots.txt setups you can copy, and the firewall rules that override them.
By InTheAnswer Editorial · Updated · 8 min read
To configure robots.txt for AI crawlers, decide by purpose rather than by company: allow the search crawlers that let engines cite you (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot), then make a separate choice about the training crawlers and tokens (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot). Each bot follows the group in robots.txt that names it, so you can allow one and block another from the same company.
The distinction matters because the cost of a mistake is lopsided. Blocking a training crawler is a statement about future models. Blocking a search crawler can remove your pages from that engine's answers, which is rarely what a marketing team intends.
Below are the crawlers sorted by purpose, three setups you can adapt, the limits of robots.txt, and the firewall settings that can override it. For how each engine uses what it crawls, see how AI answer engines choose sources.
Three kinds of AI crawler
AI companies document three kinds of automated visitor, each with a different job.
Search crawlers
These build the indexes engines retrieve from when they answer with sources. Blocking them is what takes a site out of AI answers.
- OAI-SearchBot (OpenAI): surfaces sites in ChatGPT's search features. OpenAI's crawler documentation says sites opted out of it won't be shown in ChatGPT search answers, though they can still appear as navigational links.
- Claude-SearchBot (Anthropic): crawls to improve the quality of Claude's search results. Anthropic's crawler help article says blocking it stops your content being indexed for search.
- PerplexityBot (Perplexity): surfaces and links sites in Perplexity's results. Perplexity's bot documentation says it isn't used to train AI models.
- Googlebot (Google): crawls for Google Search, including AI Overviews and AI Mode.
- Bingbot (Microsoft): crawls for Bing, which grounds Copilot's web answers.
Training crawlers and tokens
These govern whether content may be used to train models. Blocking them doesn't, by the companies' own accounts, remove a site from search.
- GPTBot (OpenAI): collects content that may be used to train OpenAI's models. OpenAI says its settings are independent, so blocking GPTBot doesn't block OAI-SearchBot.
- ClaudeBot (Anthropic): collects content that may contribute to model training.
- Google-Extended (Google): not a separate crawler but a control token. It covers training of Gemini models and grounding in Gemini Apps and Vertex AI, and Google says it doesn't affect inclusion or ranking in Google Search.
- Applebot-Extended (Apple): also a token that doesn't crawl. It opts content out of training Apple's foundation models, and Apple says disallowed pages can still appear in its search features.
- CCBot (Common Crawl): builds an open web archive whose public datasets have been used, among other things, to train language models, so blocking it is usually treated as a training decision.
Microsoft has no separate AI token. In 2023 it said the nocache and noarchive robots meta tags control how content is used in what was then Bing Chat and in training, while leaving it in Bing search results.
User-initiated fetchers
These visit a page because a person asked the assistant to, such as pasting a URL or asking a question that needs a live lookup. ChatGPT-User, Claude-User and Perplexity-User fall into this group, and robots.txt treatment varies. OpenAI says robots.txt rules may not apply to ChatGPT-User and that it isn't used to decide what appears in search. Perplexity says Perplexity-User generally ignores robots.txt. Anthropic says its bots honor robots.txt, and that blocking Claude-User may reduce your visibility in user-directed searches.
How robots.txt matching works
A few rules from the robots.txt standard (RFC 9309) explain most configuration mistakes.
- Each bot obeys one group. A crawler looks for the group whose
User-agentline matches its name, case-insensitively. Only if none matches does it fall back to theUser-agent: *group. - A named group replaces the wildcard group. If you add a group for GPTBot, GPTBot ignores everything in
*. Any private paths you want it to skip must be repeated in its own group. - The longest matching path wins.
Disallow: /checkout/beatsAllow: /for URLs under/checkout/. - One file per host.
www.example.comandshop.example.comeach need their own robots.txt at the root. - Several bots can share a group. Stacking
User-agentlines over one set of rules is standard; if you doubt a parser handles it, give each bot its own group.
Three robots.txt setups you can copy
Each example uses example.com and keeps checkout and account pages private. Adjust the paths to your own site.
Allow everything
If you want every search engine and AI crawler to read your public pages, a short file does it.
User-agent: *
Allow: /
Disallow: /checkout/
Disallow: /account/
Sitemap: https://www.example.com/sitemap.xmlEvery crawler without its own group follows these rules, so you don't need to name AI bots to allow them. Some sites name them anyway to make the intent explicit; InTheAnswer's robots.txt does.
Allow AI search, opt out of training
This suits brands that want to be cited but don't want their content used for training. Training crawlers and tokens get their own group; everything else, search crawlers included, follows the wildcard group.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
User-agent: *
Allow: /
Disallow: /checkout/
Disallow: /account/
Sitemap: https://www.example.com/sitemap.xmlTwo trade-offs are worth knowing. Blocking Google-Extended also opts you out of grounding in Gemini Apps, according to Google, so it isn't purely a training switch. And a site that blocks training crawlers may still appear in a model's background knowledge from data gathered before the block.
If your wildcard group blocks unknown bots by default, flip the approach and allow the search crawlers by name:
User-agent: *
Disallow: /
User-agent: Googlebot
User-agent: Bingbot
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
Disallow: /checkout/
Disallow: /account/Block one bot
To block a single crawler and leave everyone else alone, give it a group with Disallow: /.
User-agent: CCBot
Disallow: /To block one bot from one section, remember that its group replaces the wildcard group, so repeat your private paths:
User-agent: PerplexityBot
Disallow: /members/
Disallow: /checkout/
Disallow: /account/What robots.txt can't do
Robots.txt is a request, not a lock. The standard says its rules aren't a form of access authorization, and Google notes that while reputable crawlers obey it, others might not. Keep these limits in mind:
- User-initiated fetchers may not follow it, as OpenAI and Perplexity document for their user agents.
- It doesn't remove pages from search. A blocked URL can still be indexed if other pages link to it, and a crawler blocked by robots.txt never sees a
noindextag. Usenoindexon a crawlable page to keep it out. - User agents can be faked. OpenAI, Perplexity, Apple and Common Crawl publish IP ranges so you can check that a visitor claiming to be their bot really is.
- Google has a second switch. Search Console now includes a setting to exclude a site from AI Overviews, AI Mode and Discover's generative AI features. It's separate from robots.txt, and Google says it doesn't affect ranking elsewhere in Search. The Google AI Overviews guide covers the snippet controls that also apply.
Firewalls and CDNs that block crawlers silently
Often the reason an AI search crawler can't read a site isn't robots.txt at all. It's a bot-management rule, web application firewall or CDN setting that returns a 403, a CAPTCHA or a JavaScript challenge. Robots.txt says yes, the firewall says no, and nobody notices.
Typical culprits include one-click "block AI bots" options (Cloudflare, for example, offers AI crawler controls), aggressive rate limits, country blocks, and challenge pages served to any non-browser client. A related trap is the robots.txt file itself: under the standard, if it returns a server error, crawlers should assume the whole site is disallowed. If you allowlist crawlers in a firewall, Perplexity recommends matching both the user-agent string and its published IP ranges, which also keeps out impostors using a borrowed name.
How to test your setup
- Open
/robots.txton every host you care about, with and withoutwww, and confirm it returns a 200 status as plain text. - Read it as each bot would: find the group that names the bot, or fall back to
*, and check the paths you care about. - Use the robots.txt report in Google Search Console for Googlebot.
- Request a page with a crawler's user-agent string, for example with
curl -A, and look for 403s or challenge pages. Firewalls that verify IP addresses may treat this differently, so it's only a rough check. - Check server or CDN logs for real visits from each crawler and the status codes they received; how to measure AI citations covers log checks.
- Allow time for changes to land. OpenAI and Perplexity each say updates can take around 24 hours to reach their systems.
The free AEO audit automates the robots.txt check for any URL, crawler by crawler.
Why this matters for placement pages too
Everything above applies to the publications you appear on, not just your own site. A placement on a publication that blocks OAI-SearchBot can't be retrieved by ChatGPT search, however strong the domain looks. One that blocks only GPTBot has opted out of training but remains available for search. Firewall rules and Google's AI-features setting are invisible from outside, so if an article never shows up in any engine, ask the publisher.
That's why crawler access is one of the first checks in what makes a publication citable. When comparing placements in the catalog, check each publication's robots.txt for the search crawlers you care about.
Frequently asked questions
Will blocking GPTBot remove my site from ChatGPT?+
Not according to OpenAI. GPTBot is the training crawler, and OpenAI says each of its settings is independent. Inclusion in ChatGPT search answers depends on OAI-SearchBot, so you can block GPTBot and still be cited.
Does blocking Google-Extended keep me out of AI Overviews?+
No. Google says Google-Extended doesn't affect inclusion in Google Search, and AI Overviews draw on the Search index crawled by Googlebot. To limit AI Overviews and AI Mode, Google points to snippet controls such as nosnippet and to its Search Console setting for generative AI features.
Should I block ChatGPT-User or Perplexity-User?+
Usually not, if you want AI visibility. These fetchers act on behalf of a person asking about your content, and blocking them can stop an assistant reading your page at the moment someone asks. OpenAI and Perplexity also say robots.txt may not apply to them, so blocking them dependably takes a firewall rule.
How quickly do AI crawlers pick up robots.txt changes?+
OpenAI and Perplexity both say it can take around 24 hours for their systems to reflect a change, and the standard says crawlers shouldn't cache robots.txt for longer than a day in normal conditions. Pages already indexed may take longer to drop out or reappear.
Top Tech / SaaS placements by AEO score
How ChatGPT, Perplexity and Google AI Overviews Choose Sources to Cite
Retrieval-augmented generation in plain language, plus what each answer engine publicly documents about its crawlers and index, and where the documentation stops.
8 min readDo Backlinks Help You Appear in Google AI Overviews?
Google says AI Overviews and AI Mode link to pages from its search index. Here's how that makes backlinks matter indirectly, and why a publication that mentions you can be the cited source.
7 min readWhat Is llms.txt and Does Your Site Need One?
llms.txt is a proposed Markdown file that gives AI tools a curated map of your site. Here's the format, what it can and can't do for AI visibility, and when it's worth publishing.
8 min read