Guide · Content & technical

How to Configure robots.txt for AI Crawlers

AI companies run separate bots for search, training and user requests. Here's how to tell them apart, three robots.txt setups you can copy, and the firewall rules that override them.

By InTheAnswer Editorial · Updated · 8 min read

To configure robots.txt for AI crawlers, decide by purpose rather than by company: allow the search crawlers that let engines cite you (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot), then make a separate choice about the training crawlers and tokens (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot). Each bot follows the group in robots.txt that names it, so you can allow one and block another from the same company.

The distinction matters because the cost of a mistake is lopsided. Blocking a training crawler is a statement about future models. Blocking a search crawler can remove your pages from that engine's answers, which is rarely what a marketing team intends.

Below are the crawlers sorted by purpose, three setups you can adapt, the limits of robots.txt, and the firewall settings that can override it. For how each engine uses what it crawls, see how AI answer engines choose sources.

Three kinds of AI crawler

AI companies document three kinds of automated visitor, each with a different job.

Search crawlers

These build the indexes engines retrieve from when they answer with sources. Blocking them is what takes a site out of AI answers.

  • OAI-SearchBot (OpenAI): surfaces sites in ChatGPT's search features. OpenAI's crawler documentation says sites opted out of it won't be shown in ChatGPT search answers, though they can still appear as navigational links.
  • Claude-SearchBot (Anthropic): crawls to improve the quality of Claude's search results. Anthropic's crawler help article says blocking it stops your content being indexed for search.
  • PerplexityBot (Perplexity): surfaces and links sites in Perplexity's results. Perplexity's bot documentation says it isn't used to train AI models.
  • Googlebot (Google): crawls for Google Search, including AI Overviews and AI Mode.
  • Bingbot (Microsoft): crawls for Bing, which grounds Copilot's web answers.

Training crawlers and tokens

These govern whether content may be used to train models. Blocking them doesn't, by the companies' own accounts, remove a site from search.

  • GPTBot (OpenAI): collects content that may be used to train OpenAI's models. OpenAI says its settings are independent, so blocking GPTBot doesn't block OAI-SearchBot.
  • ClaudeBot (Anthropic): collects content that may contribute to model training.
  • Google-Extended (Google): not a separate crawler but a control token. It covers training of Gemini models and grounding in Gemini Apps and Vertex AI, and Google says it doesn't affect inclusion or ranking in Google Search.
  • Applebot-Extended (Apple): also a token that doesn't crawl. It opts content out of training Apple's foundation models, and Apple says disallowed pages can still appear in its search features.
  • CCBot (Common Crawl): builds an open web archive whose public datasets have been used, among other things, to train language models, so blocking it is usually treated as a training decision.

Microsoft has no separate AI token. In 2023 it said the nocache and noarchive robots meta tags control how content is used in what was then Bing Chat and in training, while leaving it in Bing search results.

User-initiated fetchers

These visit a page because a person asked the assistant to, such as pasting a URL or asking a question that needs a live lookup. ChatGPT-User, Claude-User and Perplexity-User fall into this group, and robots.txt treatment varies. OpenAI says robots.txt rules may not apply to ChatGPT-User and that it isn't used to decide what appears in search. Perplexity says Perplexity-User generally ignores robots.txt. Anthropic says its bots honor robots.txt, and that blocking Claude-User may reduce your visibility in user-directed searches.

How robots.txt matching works

A few rules from the robots.txt standard (RFC 9309) explain most configuration mistakes.

  • Each bot obeys one group. A crawler looks for the group whose User-agent line matches its name, case-insensitively. Only if none matches does it fall back to the User-agent: * group.
  • A named group replaces the wildcard group. If you add a group for GPTBot, GPTBot ignores everything in *. Any private paths you want it to skip must be repeated in its own group.
  • The longest matching path wins. Disallow: /checkout/ beats Allow: / for URLs under /checkout/.
  • One file per host. www.example.com and shop.example.com each need their own robots.txt at the root.
  • Several bots can share a group. Stacking User-agent lines over one set of rules is standard; if you doubt a parser handles it, give each bot its own group.

Three robots.txt setups you can copy

Each example uses example.com and keeps checkout and account pages private. Adjust the paths to your own site.

Allow everything

If you want every search engine and AI crawler to read your public pages, a short file does it.

User-agent: *
Allow: /
Disallow: /checkout/
Disallow: /account/

Sitemap: https://www.example.com/sitemap.xml

Every crawler without its own group follows these rules, so you don't need to name AI bots to allow them. Some sites name them anyway to make the intent explicit; InTheAnswer's robots.txt does.

Allow AI search, opt out of training

This suits brands that want to be cited but don't want their content used for training. Training crawlers and tokens get their own group; everything else, search crawlers included, follows the wildcard group.

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

User-agent: *
Allow: /
Disallow: /checkout/
Disallow: /account/

Sitemap: https://www.example.com/sitemap.xml

Two trade-offs are worth knowing. Blocking Google-Extended also opts you out of grounding in Gemini Apps, according to Google, so it isn't purely a training switch. And a site that blocks training crawlers may still appear in a model's background knowledge from data gathered before the block.

If your wildcard group blocks unknown bots by default, flip the approach and allow the search crawlers by name:

User-agent: *
Disallow: /

User-agent: Googlebot
User-agent: Bingbot
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
Disallow: /checkout/
Disallow: /account/

Block one bot

To block a single crawler and leave everyone else alone, give it a group with Disallow: /.

User-agent: CCBot
Disallow: /

To block one bot from one section, remember that its group replaces the wildcard group, so repeat your private paths:

User-agent: PerplexityBot
Disallow: /members/
Disallow: /checkout/
Disallow: /account/
Key takeaway: Decide per purpose, not per company. Allow the search crawlers you want citations from, decide separately about training, and remember that any bot with its own group stops reading the wildcard rules.

What robots.txt can't do

Robots.txt is a request, not a lock. The standard says its rules aren't a form of access authorization, and Google notes that while reputable crawlers obey it, others might not. Keep these limits in mind:

  • User-initiated fetchers may not follow it, as OpenAI and Perplexity document for their user agents.
  • It doesn't remove pages from search. A blocked URL can still be indexed if other pages link to it, and a crawler blocked by robots.txt never sees a noindex tag. Use noindex on a crawlable page to keep it out.
  • User agents can be faked. OpenAI, Perplexity, Apple and Common Crawl publish IP ranges so you can check that a visitor claiming to be their bot really is.
  • Google has a second switch. Search Console now includes a setting to exclude a site from AI Overviews, AI Mode and Discover's generative AI features. It's separate from robots.txt, and Google says it doesn't affect ranking elsewhere in Search. The Google AI Overviews guide covers the snippet controls that also apply.

Firewalls and CDNs that block crawlers silently

Often the reason an AI search crawler can't read a site isn't robots.txt at all. It's a bot-management rule, web application firewall or CDN setting that returns a 403, a CAPTCHA or a JavaScript challenge. Robots.txt says yes, the firewall says no, and nobody notices.

Typical culprits include one-click "block AI bots" options (Cloudflare, for example, offers AI crawler controls), aggressive rate limits, country blocks, and challenge pages served to any non-browser client. A related trap is the robots.txt file itself: under the standard, if it returns a server error, crawlers should assume the whole site is disallowed. If you allowlist crawlers in a firewall, Perplexity recommends matching both the user-agent string and its published IP ranges, which also keeps out impostors using a borrowed name.

How to test your setup

  1. Open /robots.txt on every host you care about, with and without www, and confirm it returns a 200 status as plain text.
  2. Read it as each bot would: find the group that names the bot, or fall back to *, and check the paths you care about.
  3. Use the robots.txt report in Google Search Console for Googlebot.
  4. Request a page with a crawler's user-agent string, for example with curl -A, and look for 403s or challenge pages. Firewalls that verify IP addresses may treat this differently, so it's only a rough check.
  5. Check server or CDN logs for real visits from each crawler and the status codes they received; how to measure AI citations covers log checks.
  6. Allow time for changes to land. OpenAI and Perplexity each say updates can take around 24 hours to reach their systems.

The free AEO audit automates the robots.txt check for any URL, crawler by crawler.

Why this matters for placement pages too

Everything above applies to the publications you appear on, not just your own site. A placement on a publication that blocks OAI-SearchBot can't be retrieved by ChatGPT search, however strong the domain looks. One that blocks only GPTBot has opted out of training but remains available for search. Firewall rules and Google's AI-features setting are invisible from outside, so if an article never shows up in any engine, ask the publisher.

That's why crawler access is one of the first checks in what makes a publication citable. When comparing placements in the catalog, check each publication's robots.txt for the search crawlers you care about.

Frequently asked questions

Will blocking GPTBot remove my site from ChatGPT?+

Not according to OpenAI. GPTBot is the training crawler, and OpenAI says each of its settings is independent. Inclusion in ChatGPT search answers depends on OAI-SearchBot, so you can block GPTBot and still be cited.

Does blocking Google-Extended keep me out of AI Overviews?+

No. Google says Google-Extended doesn't affect inclusion in Google Search, and AI Overviews draw on the Search index crawled by Googlebot. To limit AI Overviews and AI Mode, Google points to snippet controls such as nosnippet and to its Search Console setting for generative AI features.

Should I block ChatGPT-User or Perplexity-User?+

Usually not, if you want AI visibility. These fetchers act on behalf of a person asking about your content, and blocking them can stop an assistant reading your page at the moment someone asks. OpenAI and Perplexity also say robots.txt may not apply to them, so blocking them dependably takes a firewall rule.

How quickly do AI crawlers pick up robots.txt changes?+

OpenAI and Perplexity both say it can take around 24 hours for their systems to reflect a change, and the standard says crawlers shouldn't cache robots.txt for longer than a day in normal conditions. Pages already indexed may take longer to drop out or reappear.

From the catalog

Top Tech / SaaS placements by AEO score

All Tech / SaaS links →
EliteAEO 96.0 Verified
Tech / SaaS
DR
97
Ahrefs
DA
96
Moz
Traffic
46M
Ahrefs
United States 24%·Nofollow
$489/ placement
EliteAEO 89.8 Verified
Tech / SaaS
DR
93
Ahrefs
DA
—
Moz
Traffic
4.4M
Ahrefs
United States 46%
$2,100/ placement
EliteAEO 88.0 Verified
Business / FinanceTech / SaaS
DR
92
Ahrefs
DA
94
Moz
Traffic
5.7M
Ahrefs
India 93%·Nofollow
$1,250/ placement
Keep reading