AI crawlers and indexing bots: the reference profiles
GPTBot, ClaudeBot, PerplexityBot, Googlebot… every robot visiting your site has an operator, a purpose and rules of its own. These reference profiles document, for each crawler, the verifiable facts: the user-agent string, the robots.txt token, the official IP ranges when the operator publishes them, and what its visits mean for your SEO and your visibility in AI answers (GEO).
Every technical fact links to its primary source — the operator's official documentation. The IP ranges shown are refreshed automatically on every site deployment, from the files published by the operators themselves.
OpenAI — ChatGPT
GPTBot
Training crawler
GPTBot is OpenAI's crawler collecting content to train its models. User-agent string, official IP ranges, how to verify it and how to block it.
OAI-SearchBot
ChatGPT search index
OAI-SearchBot builds the index that surfaces sites in ChatGPT search. User-agent, IP ranges, GEO stakes: blocking it removes you from the results.
ChatGPT-User
Live browsing (user-triggered)
ChatGPT-User fetches a page when a ChatGPT user asks for it. Not a bulk crawler: every hit is a potential reader. IP ranges, robots.txt, measurement.
Anthropic — Claude
ClaudeBot
Training crawler
ClaudeBot collects web content to train Anthropic's Claude models. User-agent, robots.txt compliance, and why IP verification of this bot is limited.
Claude-SearchBot
Search result quality
Claude-SearchBot improves the quality of Claude's search results. Role, user-agent, robots.txt: the Anthropic bot to allow for visibility.
Claude-User
Live browsing (user-triggered)
Claude-User fetches a page when a Claude user asks for it. Every hit is a real human read through AI. Detection, robots.txt, measurement in your logs.
Google — Search, AI Overviews & Gemini
Google-Extended
robots.txt token (Gemini training)
Google-Extended is not a crawler: it's a robots.txt token controlling whether your content trains Gemini — without touching your SEO or rankings.
GoogleOther
Generic crawler (R&D, products)
GoogleOther is Google's “generic” crawler, used outside Search indexing: R&D, products, internal uses. Detection, IP ranges, robots.txt token.
Googlebot
Google Search indexing
Googlebot feeds the Google Search index — and AI Overviews. User-agent, official IP ranges, anti-spoofing verification, robots.txt best practices.
Perplexity — answer engine
Microsoft — Bing & Copilot
Apple — Siri & Apple Intelligence
To find out which of these bots actually visit your site — and verify their authenticity by IP — drop your server logs into our analyzer: everything happens in your browser, no upload.
Analyze my logs