Mover Marketing AI

What are AI crawlers?

AI crawlers are bots that AI companies use to collect web pages for model training, to build AI search results, or to fetch a page when a user asks, and site owners control most of them separately through robots.txt.

ExplainerReviewed by Nicholas DiMoro

Learning objectives

After reading this article you will be able to:

  • Tell apart AI crawlers used for training, search, and user requests
  • Identify the main AI crawlers by company and user agent
  • Decide which AI crawlers a moving company should allow

What are AI crawlers?

AI crawlers are bots that AI companies use to read web pages. They do one of three jobs:

  • Training: collect content that may be used to train AI models.
  • Search: index pages so an AI search tool can show and cite them.
  • User requests: fetch a page in real time because a user asked a question.

Cloudflare uses similar groups for its bot settings: "Search," "Agent" (activity "acting in real time on a person's behalf"), and "Training."

The difference matters because blocking one kind doesn't block the others.

Which AI crawlers are most common?

The main AI companies publish their crawler names, also called user agents:

CompanySearchTrainingUser requests
OpenAIOAI-SearchBotGPTBotChatGPT-User
AnthropicClaude-SearchBotClaudeBotClaude-User
PerplexityPerplexityBotNot listedPerplexity-User
GoogleGooglebotGoogle-Extended (a control token)Not listed here
AppleApplebotApplebot-Extended (a control token)Not listed here

What each company says about them:

  • OpenAI. OAI-SearchBot is used "to surface websites in search results in ChatGPT's search features." GPTBot crawls "content that may be used in training our generative AI foundation models." OpenAI says each setting "is independent of the others."
  • Anthropic. ClaudeBot collects content "that could potentially contribute to" model training. Blocking Claude-SearchBot "may reduce your site's visibility and accuracy in user search results."
  • Perplexity. PerplexityBot "is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models."
  • Google. Google-Extended has no crawler of its own; it's a robots.txt token that controls whether content Google crawls may be used for training Gemini models and for grounding in Gemini Apps. Google says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."
  • Apple. Applebot powers search in Spotlight, Siri, and Safari. Publishers can opt out of generative model training "by disallowing Applebot-Extended in the robots.txt file."

Do AI crawlers follow robots.txt?

The companies say their automated crawlers do. Google, for example, says its common crawlers "always respect robots.txt rules for automatic crawls." User-triggered fetchers vary:

  • OpenAI says that because ChatGPT-User visits are "initiated by a user, robots.txt rules may not apply."
  • Perplexity says Perplexity-User "generally ignores robots.txt rules."
  • Anthropic says its bots "respect 'do not crawl' signals by honoring industry standard directives in robots.txt."

Robots.txt controls crawling, not what's already known. Google notes that robots.txt "is not a mechanism for keeping a web page out of Google."

How do you allow or block a specific AI crawler?

Add a group for its user agent to the site's robots.txt file. For example, to keep a site out of OpenAI's training while staying in ChatGPT search:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

Anthropic notes that blocking by IP address instead "may not work correctly or persistently guarantee an opt-out," because it stops the crawler from reading robots.txt.

Can a firewall block AI crawlers by accident?

Yes. CDN and security services can block bots regardless of robots.txt. Cloudflare lets customers block AI bots by category. Under defaults it announced for new domains starting September 15, 2026, "bots classified as Training or as Agent will be blocked on pages that display ads, and Search will remain allowed."

A site that wants to appear in AI search answers should check both robots.txt and its CDN or firewall settings.

Which AI crawlers should a moving company allow?

That depends on the business's goals, but the search crawlers are the ones that affect visibility. A mover that wants to be cited in ChatGPT, Claude, or Perplexity answers needs OAI-SearchBot, Claude-SearchBot, and PerplexityBot to reach its site. Training crawlers are a separate decision about whether content may be used to train models. See how businesses get recommended by ChatGPT.

FAQs

What is an AI crawler?
An AI crawler is a bot that an AI company uses to read web pages. Some collect content to train AI models, some index pages so an AI search tool can cite them, and some fetch a page when a user asks a question.
Is ChatGPT a web crawler?
ChatGPT itself isn't, but OpenAI runs crawlers for it. OAI-SearchBot surfaces websites in ChatGPT's search features, GPTBot collects content that may be used for training, and ChatGPT-User visits pages when a user's request needs them.
Does blocking AI training crawlers remove a site from AI search?
Not for the companies that separate them. OpenAI, Anthropic, and Perplexity each use a different crawler for search than for training, and Google says blocking Google-Extended doesn't affect a site's inclusion in Google Search.
Can a firewall block AI crawlers even if robots.txt allows them?
Yes. CDN and firewall services can block bots on their own. Cloudflare, for example, lets customers block AI bots by category, and new defaults block training and agent bots on pages with ads.

Our take

Sources

Want this handled for your moving company?

Explore SEO for movers