What are AI crawlers?
AI crawlers are bots that AI companies use to read web pages. They do one of three jobs:
- Training: collect content that may be used to train AI models.
- Search: index pages so an AI search tool can show and cite them.
- User requests: fetch a page in real time because a user asked a question.
Cloudflare uses similar groups for its bot settings: "Search," "Agent" (activity "acting in real time on a person's behalf"), and "Training."
The difference matters because blocking one kind doesn't block the others.
Which AI crawlers are most common?
The main AI companies publish their crawler names, also called user agents:
| Company | Search | Training | User requests |
|---|---|---|---|
| OpenAI | OAI-SearchBot | GPTBot | ChatGPT-User |
| Anthropic | Claude-SearchBot | ClaudeBot | Claude-User |
| Perplexity | PerplexityBot | Not listed | Perplexity-User |
| Googlebot | Google-Extended (a control token) | Not listed here | |
| Apple | Applebot | Applebot-Extended (a control token) | Not listed here |
What each company says about them:
- OpenAI. OAI-SearchBot is used "to surface websites in search results in ChatGPT's search features." GPTBot crawls "content that may be used in training our generative AI foundation models." OpenAI says each setting "is independent of the others."
- Anthropic. ClaudeBot collects content "that could potentially contribute to" model training. Blocking Claude-SearchBot "may reduce your site's visibility and accuracy in user search results."
- Perplexity. PerplexityBot "is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models."
- Google. Google-Extended has no crawler of its own; it's a robots.txt token that controls whether content Google crawls may be used for training Gemini models and for grounding in Gemini Apps. Google says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."
- Apple. Applebot powers search in Spotlight, Siri, and Safari. Publishers can opt out of generative model training "by disallowing Applebot-Extended in the robots.txt file."
Do AI crawlers follow robots.txt?
The companies say their automated crawlers do. Google, for example, says its common crawlers "always respect robots.txt rules for automatic crawls." User-triggered fetchers vary:
- OpenAI says that because ChatGPT-User visits are "initiated by a user, robots.txt rules may not apply."
- Perplexity says Perplexity-User "generally ignores robots.txt rules."
- Anthropic says its bots "respect 'do not crawl' signals by honoring industry standard directives in robots.txt."
Robots.txt controls crawling, not what's already known. Google notes that robots.txt "is not a mechanism for keeping a web page out of Google."
How do you allow or block a specific AI crawler?
Add a group for its user agent to the site's robots.txt file. For example, to keep a site out of OpenAI's training while staying in ChatGPT search:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
Anthropic notes that blocking by IP address instead "may not work correctly or persistently guarantee an opt-out," because it stops the crawler from reading robots.txt.
Can a firewall block AI crawlers by accident?
Yes. CDN and security services can block bots regardless of robots.txt. Cloudflare lets customers block AI bots by category. Under defaults it announced for new domains starting September 15, 2026, "bots classified as Training or as Agent will be blocked on pages that display ads, and Search will remain allowed."
A site that wants to appear in AI search answers should check both robots.txt and its CDN or firewall settings.
Which AI crawlers should a moving company allow?
That depends on the business's goals, but the search crawlers are the ones that affect visibility. A mover that wants to be cited in ChatGPT, Claude, or Perplexity answers needs OAI-SearchBot, Claude-SearchBot, and PerplexityBot to reach its site. Training crawlers are a separate decision about whether content may be used to train models. See how businesses get recommended by ChatGPT.