Skip to main content
Pen and ink illustration of a harbour gate with a lantern, letting some small boats through and holding others back

SEO

Block AI Crawlers? GPTBot and robots.txt Explained

Blocking AI crawlers is not one decision. Each provider runs different bots for different jobs. Here is what each does and how robots.txt controls them.

Written by Craig Fearn

Director

Last updated: 2 October 2026

Part of our complete guide: AI SEO: What It Is and What Actually Works

Blocking AI crawlers is not one decision. OpenAI, Perplexity and Anthropic each publish several bots with different jobs, and you can allow some and block others in robots.txt. If you block the bots that power search, the documentation says you shouldn’t expect to be shown by that product.

This post sets out what each provider says its bots do, how robots.txt controls them, and what ours looks like. It sits under our AI SEO guide. Our view throughout is labelled as such, and we don’t promise any result.

The bots, provider by provider

OpenAI (bot documentation):

  • OAI-SearchBot surfaces sites in ChatGPT search. A site that opts out won’t be shown in ChatGPT search answers.
  • GPTBot is for training.
  • ChatGPT-User acts on user requests, and robots.txt may not apply to it.

Perplexity (bot documentation):

  • PerplexityBot surfaces and links sites in results, isn’t used for training and obeys robots.txt.
  • Perplexity-User generally ignores robots.txt.

Anthropic (crawler help article):

  • ClaudeBot is for training.
  • Claude-User handles user requests.
  • Claude-SearchBot relates to search quality.
  • All three can be controlled via robots.txt.

Google’s position is different. Its AI features documentation says there are no extra requirements to appear in AI Overviews or AI Mode, and that the snippet controls apply to them. We cover those below.

How robots.txt controls them

If you have never opened yours, type your domain followed by /robots.txt into a browser.

robots.txt is a plain text file at the root of your site. Each group starts with a User-agent line naming a bot, followed by Allow or Disallow rules. OpenAI says each of its bots is controlled separately, so a rule for one does not change another.

A bot with no group of its own follows the general User-agent: * group. That is standard robots.txt behaviour, and it means a site with an open general rule is already allowing bots it never names.

Bots that may ignore robots.txt

This part is easy to miss, and it changes what robots.txt can promise you.

Read the documentation here, because it is specific. OpenAI says robots.txt may not apply to ChatGPT-User. Perplexity says Perplexity-User generally ignores robots.txt. Both bots act on a user’s request rather than crawling on their own.

So robots.txt is not a lock. If something must stay private, don’t publish it. Also notice the pattern in the documentation: the bots described as ignoring or possibly ignoring robots.txt are the ones that act on a user’s request, while PerplexityBot is described as obeying it.

The trade-off of blocking search bots

Many site owners want to stay out of training while staying visible in search. The documentation supports separating those:

  • GPTBot and ClaudeBot are described as training bots.
  • OAI-SearchBot and PerplexityBot are described as the ones that surface sites in their search results.

If you block a search bot, you shouldn’t expect to be shown by that product. OpenAI says so directly for ChatGPT search. Our view: unless you have a specific reason, leave the search bots allowed, and make the training decision on its own merits.

An example robots.txt

This is an example, not advice for every site. It blocks the two bots described as training bots and allows the rest through the general rule. It uses only the bot names from the documentation above.

# Example only. Decide bot by bot for your own site.
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: *
Allow: /

Keep the file simple. Every extra group is one more thing to get wrong, and a stray Disallow: / under the general User-agent: * group would block every crawler that has no group of its own, including search bots you may want. Test any change before you rely on it, and check the provider’s documentation again, because bot names and behaviour can change.

Google snippet controls

Google’s AI features documentation says nosnippet, data-nosnippet, max-snippet and noindex apply to AI Overviews and AI Mode. If you use them, you are also limiting what AI features can show from your pages.

Use them deliberately. Google also says AI feature traffic is counted under “Web” in Search Console, with no separate AI report.

What our own robots.txt says

Our file sets Allow: / for everything through the general rule, and also names GPTBot, ChatGPT-User, PerplexityBot, ClaudeBot, anthropic-ai and Google-Extended with Allow: /. It then points to our sitemap. The other bots mentioned above, such as OAI-SearchBot and Claude-SearchBot, are allowed via the general rule.

The sample above is shorter than most real files, and it leaves out anything specific to your own site, such as a sitemap line. Staying open is a choice that may suit a business that wants to be found and may not suit a publisher or a paywalled site. As our test in the AI SEO guide shows, being allowed doesn’t mean being named: we appeared in 4 of 40 answers.

What to do next

  1. Open your robots.txt and list which bots it names.
  2. Decide separately on training bots and search bots.
  3. Check your snippet controls aren’t limiting more than you intend.
  4. Read our post on generative engine optimisation for what to do beyond crawler access.

If you want help with this, see our AI SEO agency page, or AI SEO in Cornwall for local businesses. We offer no guarantee or refund, because nobody can promise what an AI engine will show.

Frequently Asked Questions

Should I block GPTBot?
It depends on what you want. OpenAI says GPTBot is for training, and each OpenAI bot is controlled separately in robots.txt (OpenAI bots). Blocking GPTBot is a separate choice from blocking OAI-SearchBot, which surfaces sites in ChatGPT search.
Does blocking GPTBot hide me from ChatGPT search?
According to OpenAI, no, because they are different bots. OAI-SearchBot surfaces sites in ChatGPT search, and a site that opts out of it won't be shown in ChatGPT search answers (OpenAI bots).
Does robots.txt stop ChatGPT-User or Perplexity-User?
Not necessarily. OpenAI says robots.txt may not apply to ChatGPT-User (OpenAI bots), and Perplexity says Perplexity-User generally ignores robots.txt (Perplexity bots).
How do I stop AI features showing my text in Google?
Google says nosnippet, data-nosnippet, max-snippet and noindex apply to AI Overviews and AI Mode, so using them also limits what those features can show from your pages (Google AI features).
Is robots.txt enough to keep content private?
No. It is a set of instructions that crawlers follow or don't, and the Perplexity and OpenAI documentation above say some user-triggered bots may not follow it. If content must stay private, don't publish it openly.

Need help with this?

We work with Cornwall small businesses on the exact challenges covered in this article. Free 30-minute call, no pressure.

Get in touch

Craig Fearn

Director

Craig is Director of Outcome Digital Marketing. He brings over a decade of C-suite advisory experience, having advised senior executives and boards on organisational strategy before focusing on the marketing decisions that move the needle for smaller businesses. As a Fellow of the Royal Society for Public Health (FRSPH) and Fellow of the Chartered Management Institute (FCMI), he applies evidence-based thinking to marketing - helping Cornwall and UK businesses make informed decisions backed by research, not hype.

Knowledge Is Only Half the Battle

You've got the insights. Now you need someone who'll help you act on them - without the runaround.

Let's Talk Strategy