Skip to content
Reference · 9 crawler identities

Nine tokens. Three purposes. Know which is which.

AI robots.txt tokens are not interchangeable. A training crawler and a search crawler control different things — this reference keeps them separated so you never block search when you only meant to opt out of training.

Check all 9 on your domain

Search and discovery

These surface sites in search and AI search experiences. Blocking them removes you from the surface they feed.

Googlebot

Google

Google Search and AI Overviews eligibility. Required for both.

OAI-SearchBot

OpenAI

Builds the ChatGPT search index. Required to appear in ChatGPT search results.

PerplexityBot

Perplexity

Builds Perplexity's search index. Required for Perplexity answers and citations.

Claude-SearchBot

Anthropic

Powers Claude web search. Blocking it can reduce visibility in Claude searches.

Bingbot

Microsoft

Bing index and Copilot grounding.

Model training and other AI use

Associated with training data or other AI product uses. Blocking these does not, by itself, block search discovery.

GPTBot

OpenAI

Crawls content that may be used for model training. Separate from OAI-SearchBot.

Google-Extended

Google

A control token (not a separate bot) for Gemini training use. Does not affect search ranking.

ClaudeBot

Anthropic

Crawls content for Claude training and retrieval. Separate from Claude-SearchBot.

Ad landing-page validation

Fetches pages to validate ad landing pages. Not an organic visibility signal.

OAI-AdsBot

OpenAI

Validates ChatGPT Ads landing pages. Relevant only if you run ChatGPT Ads.

Honesty box

Permission is not proof of a visit, indexing, ingestion, ranking, or citation. robots.txt is advisory — a firewall, CDN rule, or bot-management layer can silently block a crawler that robots.txt allows. Broader utilities may test additional bots as a separate experiment; the paid Optimus scope stays at these 9 identities.

Last updated: September 10, 2026