NameWhat it doesrobots.txtIf you block it
GooglebotGoogle · Search engineBuilds Google Search, including AI Overviews and AI Mode.Follows robots.txtYou disappear from Google Search. Almost never what you want.
Google-ExtendedGoogle · Trains AINot a crawler: a robots.txt setting for whether Gemini may train on your pages or use them in Gemini apps.robots.txt setting onlyNo effect on Google Search or ranking. It doesn't remove you from AI Overviews either.
Google-AgentGoogle · Acts for a userAgents on Google's servers that browse and take actions when a user asks. Some requests are signed.May ignore robots.txtGoogle says user-triggered agents generally ignore robots.txt. Block at your firewall if you must.
BingbotMicrosoft · Search engineBuilds Bing, which also grounds Microsoft Copilot's answers.Follows robots.txtYou disappear from Bing and from Copilot's web answers. There's no separate Copilot crawler to block.
GPTBotOpenAI · Trains AICollects pages that may be used to train OpenAI's models.Follows robots.txtYour pages aren't used for training. No effect on ChatGPT search.
OAI-SearchBotOpenAI · AI searchFinds pages to show and link in ChatGPT's search answers.Follows robots.txtOpenAI says blocked sites aren't shown in ChatGPT search answers.
ChatGPT-UserOpenAI · A user's requestFetches a page when someone asks ChatGPT about it, or a GPT needs it.May ignore robots.txtOpenAI says robots.txt rules may not apply, because a user started the request.
ChatGPT cloud browserOpenAI · Acts for a userThe browser ChatGPT uses to do tasks for someone, such as filling in a form. Signs every request as chatgpt.com.May ignore robots.txtIt isn't a robots.txt crawler. OpenAI says each site decides, and explains how to allow it in common firewalls.
ClaudeBotAnthropic · Trains AICollects pages that may be used to train Anthropic's models.Follows robots.txtFuture pages are left out of training. Anthropic also supports Crawl-delay.
Claude-SearchBotAnthropic · AI searchIndexes pages to improve Claude's search results.Follows robots.txtYour pages aren't indexed for Claude's search.
Claude-UserAnthropic · A user's requestFetches a page when someone asks Claude a question.Follows robots.txtAnthropic says this can reduce your visibility when people search through Claude. It does follow robots.txt.
PerplexityBotPerplexity · AI searchFinds pages to show and link in Perplexity's answers. Not used for model training.Follows robots.txtYou don't appear in Perplexity's search results.
Perplexity-UserPerplexity · A user's requestFetches a page when a user asks Perplexity something.May ignore robots.txtPerplexity says it generally ignores robots.txt rules.
Meta-ExternalAgentMeta · Trains AICollects pages to train Meta's AI models or improve its products.Follows robots.txtYour pages aren't collected for training.
Meta-ExternalFetcherMeta · A user's requestFetches a link when a user asks, including to help Meta's AI complete tasks.May ignore robots.txtMeta says it may bypass robots.txt rules.
ApplebotApple · Search enginePowers Siri, Spotlight and Safari search. Follows your Googlebot rules if you don't name it.Follows robots.txtYou drop out of Siri and Spotlight suggestions.
Applebot-ExtendedApple · Trains AINot a crawler: a robots.txt setting for whether Apple may train its models on your pages.robots.txt setting onlyNo effect on Apple search results or ranking.
AmazonbotAmazon · Trains AICrawls pages that may be used to train Amazon's AI models.Follows robots.txtYour pages aren't collected. Amazon also reads the noarchive tag as “don't train”.
Amzn-SearchBotAmazon · AI searchSearch experiences such as Alexa. Not used for training.Follows robots.txtYou may not appear in Alexa's answers.
Amzn-UserAmazon · A user's requestFetches pages live for a customer's question.May ignore robots.txtAmazon says it may not follow all robots.txt rules.
CCBotCommon Crawl · Trains AIBuilds a free public archive of the web, widely used to train AI models.Follows robots.txtYou're left out of future archives. Beware fake CCBots; check the IP.
21 names from the companies' own documentation, checked 29 September 2026. Agents that run inside a person's own browser, such as Claude in Chrome or Gemini in Chrome, have no documented name: to your website they look like that person's browser.