Insights · Website owners

Which AI agents should you let into your website?

More than half of the visitors to a typical website aren't people. Some collect your pages to train AI, some find them for AI search, and some are a customer's assistant acting right now. This guide shows who they are, what each one costs or brings you, and how to decide who gets in.

Visitors to a clinic's website · illustration
  1. 09:14:03Safari on iPhonePerson /prices Let inA customer
  2. 09:14:08GooglebotSearch engine /treatments Let inGoogle Search
1 people, 1 machines.Each machine wants something different. The question is which ones you let in.

Checked on 29 September 2026 against each company's own crawler documentation. Bot names and rules change often, and the clinic in the illustration is fictional. General information, not legal or security advice.

The short version

Four things to know first

“Block AI” sounds like one switch. It's really three different decisions, and one of them can cost you customers.

Not one kind of AI visitor, but three

Some collect pages to train AI, some find pages for AI search, and some act for a person right now. Each deserves a different answer.

Blocking training costs little traffic

Training crawlers send almost no visitors back. Whether to let them in is a choice about your content, not your sales.

Blocking search and agents can cost customers

AI search bots and agents are how people find, check and book you. Turn them away and those people go elsewhere.

robots.txt asks, it doesn't lock

Several bots that act for a person say they may ignore it. Real control happens at your host or network.

53%
of all web traffic in 2025 came from bots, not people
Imperva Bad Bot Report, Apr 2026
52%
of crawler requests were for AI training in June 2026, up from 22% in spring 2025
Cloudflare, Jul 2026
887 : 1
pages OpenAI's training crawler took for every visit OpenAI sent back, in early August 2025
Cloudflare, Aug 2025
+60%
higher conversion from visitors sent by AI assistants than from other traffic, in US retail
Adobe, Jul 2026
Who's knocking

Four kinds of AI visitor

Three that say who they are, and one that doesn't. What each one gives you, what turning it away costs, and where we'd start.

Your call

Trains AI

GPTBot, ClaudeBot, Google-Extended, Meta-ExternalAgent, CCBot

You get: Your pages may shape what future AI models know. Almost no visitors come back from it.

If you block it: Blocking costs very little traffic. It won't remove what was already collected.
Let in

Finds pages for AI search

OAI-SearchBot, Claude-SearchBot, PerplexityBot, and Googlebot and Bingbot for Google's and Copilot's answers

You get: Links to you in AI answers, and the visitors that come with them.

If you block it: Blocking leaves you out of those answers. OpenAI says so directly for ChatGPT search.
Let in, with limits

Acts for a person now

ChatGPT-User, Claude-User, Perplexity-User, Google-Agent, ChatGPT's cloud browser

You get: A real person asking about you, comparing you or booking you through their assistant.

If you block it: Blocking turns that person away. Several of these say they may ignore robots.txt anyway.
Slow down or block

Hides who it is

Scrapers that pretend to be a normal browser and switch addresses

You get: Load on your server and higher hosting bills. Nothing you can count.

If you block it: robots.txt won't stop them. Rate limits and bot management will, some of the time.
The directory

Every AI visitor, by name

The names the big AI companies use, what each one is for, and what blocking it actually does. Paste a line from your server logs to see who it is.

NameWhat it doesrobots.txtIf you block it
GooglebotGoogle · Search engineBuilds Google Search, including AI Overviews and AI Mode.Follows robots.txtYou disappear from Google Search. Almost never what you want.
Google-ExtendedGoogle · Trains AINot a crawler: a robots.txt setting for whether Gemini may train on your pages or use them in Gemini apps.robots.txt setting onlyNo effect on Google Search or ranking. It doesn't remove you from AI Overviews either.
Google-AgentGoogle · Acts for a userAgents on Google's servers that browse and take actions when a user asks. Some requests are signed.May ignore robots.txtGoogle says user-triggered agents generally ignore robots.txt. Block at your firewall if you must.
BingbotMicrosoft · Search engineBuilds Bing, which also grounds Microsoft Copilot's answers.Follows robots.txtYou disappear from Bing and from Copilot's web answers. There's no separate Copilot crawler to block.
GPTBotOpenAI · Trains AICollects pages that may be used to train OpenAI's models.Follows robots.txtYour pages aren't used for training. No effect on ChatGPT search.
OAI-SearchBotOpenAI · AI searchFinds pages to show and link in ChatGPT's search answers.Follows robots.txtOpenAI says blocked sites aren't shown in ChatGPT search answers.
ChatGPT-UserOpenAI · A user's requestFetches a page when someone asks ChatGPT about it, or a GPT needs it.May ignore robots.txtOpenAI says robots.txt rules may not apply, because a user started the request.
ChatGPT cloud browserOpenAI · Acts for a userThe browser ChatGPT uses to do tasks for someone, such as filling in a form. Signs every request as chatgpt.com.May ignore robots.txtIt isn't a robots.txt crawler. OpenAI says each site decides, and explains how to allow it in common firewalls.
ClaudeBotAnthropic · Trains AICollects pages that may be used to train Anthropic's models.Follows robots.txtFuture pages are left out of training. Anthropic also supports Crawl-delay.
Claude-SearchBotAnthropic · AI searchIndexes pages to improve Claude's search results.Follows robots.txtYour pages aren't indexed for Claude's search.
Claude-UserAnthropic · A user's requestFetches a page when someone asks Claude a question.Follows robots.txtAnthropic says this can reduce your visibility when people search through Claude. It does follow robots.txt.
PerplexityBotPerplexity · AI searchFinds pages to show and link in Perplexity's answers. Not used for model training.Follows robots.txtYou don't appear in Perplexity's search results.
Perplexity-UserPerplexity · A user's requestFetches a page when a user asks Perplexity something.May ignore robots.txtPerplexity says it generally ignores robots.txt rules.
Meta-ExternalAgentMeta · Trains AICollects pages to train Meta's AI models or improve its products.Follows robots.txtYour pages aren't collected for training.
Meta-ExternalFetcherMeta · A user's requestFetches a link when a user asks, including to help Meta's AI complete tasks.May ignore robots.txtMeta says it may bypass robots.txt rules.
ApplebotApple · Search enginePowers Siri, Spotlight and Safari search. Follows your Googlebot rules if you don't name it.Follows robots.txtYou drop out of Siri and Spotlight suggestions.
Applebot-ExtendedApple · Trains AINot a crawler: a robots.txt setting for whether Apple may train its models on your pages.robots.txt setting onlyNo effect on Apple search results or ranking.
AmazonbotAmazon · Trains AICrawls pages that may be used to train Amazon's AI models.Follows robots.txtYour pages aren't collected. Amazon also reads the noarchive tag as “don't train”.
Amzn-SearchBotAmazon · AI searchSearch experiences such as Alexa. Not used for training.Follows robots.txtYou may not appear in Alexa's answers.
Amzn-UserAmazon · A user's requestFetches pages live for a customer's question.May ignore robots.txtAmazon says it may not follow all robots.txt rules.
CCBotCommon Crawl · Trains AIBuilds a free public archive of the web, widely used to train AI models.Follows robots.txtYou're left out of future archives. Beware fake CCBots; check the IP.

21 names from the companies' own documentation, checked 29 September 2026. Agents that run inside a person's own browser, such as Claude in Chrome or Gemini in Chrome, have no documented name: to your website they look like that person's browser.

You're the bouncer

Decide who gets in

Eight visitors at your door. Let each one in or turn it away, see what happens, and copy the robots.txt your choices build.

Search engineGooglebotGoogle

Wants to read your pages for Google Search.

robots.txt
# Built on sproutmedia.ie. Check before publishing.

User-agent: *
Allow: /
Beyond robots.txt

Five layers of control

robots.txt is only the first layer. Pick a layer to see what it can stop, what it can't, and what to watch out for.

robots.txtA text file at yoursite.com/robots.txt, or a setting in Squarespace, Wix or WordPress.com. Effort: 10 minutes.
  • AI training crawlers Stops or controls it

    The big AI companies say their training crawlers follow it.

  • Fetches a user asked for Partly

    OpenAI, Perplexity, Meta, Amazon and Google say some of their user-requested fetchers may ignore it.

  • Agents in a person's own browser No

    They never read it. They're just a browser.

  • Bots that hide who they are No

    It's voluntary. The standard itself says it isn't security.

Changes take time: OpenAI says about 24 hours, and Amazon may use a copy up to 30 days old.

Three traps

Where good intentions backfire

The most common ways a business blocks the wrong thing, or thinks it has blocked something it hasn't.

Trap 1Blocking AI training can block Google

Cloudflare now counts Googlebot, Bingbot and Applebot as multi-purpose crawlers, because they serve search and AI. Sites that block training there can block them too. Since July 2026, site owners have reported Google's verified crawler being refused with “Blocked by Block AI training crawlers”.

In Cloudflare, check that multi-purpose crawlers stay allowed. Then test a page with Search Console's URL Inspection.

Trap 2Google-Extended doesn't take you out of AI Overviews

Blocking Google-Extended stops Gemini training and use in Gemini apps. It doesn't change Google Search, including AI Overviews and AI Mode, which are built from Googlebot's crawl.

To leave AI Overviews and AI Mode, use the generative AI control in Search Console, worldwide since 31 August 2026.

Trap 3Blocking now doesn't delete what was taken

robots.txt and platform switches work from now on. Anthropic talks about excluding your “future materials”, and Squarespace says its switch isn't retroactive.

Decide early, and don't expect a block today to change what an AI already knows about you.

When the visitor is a customer

Agents acting for someone

The newest visitors are assistants doing a task for a person: comparing, filling in forms, booking. Some say who they are. Many can't be told apart from the person.

AgentWhere it runsWhat it does on your siteCan your site tell?
ChatGPT cloud browserOpenAI's serversBrowses and fills in forms. Pauses for sign-ins and asks the person before booking or paying.Yes. It signs every request as chatgpt.com.
Google-AgentGoogle's serversBrowses and takes actions when a user asks.Mostly. It has a name and published addresses, and some requests are signed.
Gemini in ChromeThe person's own ChromeBrowses with the person's own sign-ins. Asks before purchases.No identifier is documented.
Claude in ChromeThe person's own ChromeBrowses for the person. Won't solve CAPTCHAs or make purchases.No identifier is documented.
Copilot in EdgeThe person's own EdgeRuns in the browser. Asks the person to supervise purchases and reservations.No identifier is documented.
Meta MuseA browser in Meta's cloudOpens a browser and fills in forms. Checks before purchases.Disputed. Amazon says it hid its identity; Meta hasn't said how it identifies.
Perplexity CometThe person's own browserBrowses and shops on the person's instructions.Disputed in court. Amazon says it didn't identify itself as an agent.

What this means for you: you can't reliably block every agent, and for the pages customers need, you probably don't want to. Decide per page instead. Keep prices, hours and booking open; protect accounts, checkout and forms with rate limits and the fraud checks you already use, not with a blanket ban on bots. What an agent needs to actually book you is in our article on what happens when a customer's AI agent tries to book you.

What's happened

Six cases worth knowing

Crawlers that cost real money, a small business knocked offline, and the fights over what agents may do.

Bandwidth

Read the Docs paid for one crawler

One crawler downloaded 73 TB in May 2024, about 10 TB in a single day, costing over $5,000. Blocking AI crawlers cut the site's bandwidth by 75%.

Source: Read the Docs, Jul 2024
Outage

A seven-person shop went down

Triplegangers' product site crashed under OpenAI's training crawler, which used around 600 addresses. It had no rule for GPTBot; its terms of service didn't stop it. robots.txt rules and Cloudflare did.

Source: TechCrunch, Jan 2025
Disguise

Crawlers that pretend to be people

SourceHut's founder described crawlers using random browser names and home internet addresses, causing dozens of brief outages a week.

Source: SourceHut, Mar 2025
Dispute

Cloudflare and Perplexity disagree

Cloudflare said Perplexity used undeclared crawlers that ignored robots.txt. Perplexity said the traffic came from a browser service it uses, acting for users. It remains unresolved.

Source: Cloudflare and Perplexity, Aug 2025
Court

Amazon v. Perplexity

A US court stopped Perplexity's Comet browser shopping on Amazon in March 2026. The appeals court lifted that order on 4 August 2026, finding the user, not Perplexity, accessed Amazon. Other claims continue.

Source: Law firm summaries, Aug 2026
Blocked

Amazon shuts out Meta's Muse

Amazon started showing Muse users a notice that access by an “unauthorized AI agent” breaks its rules. Amazon says Muse hid its identity; Meta disputes how it works.

Source: TechCrunch, 21 Sep 2026
On your platform

Where to change it

Where the setting lives on the most common website platforms, what it really does, and the catch.

Where
Settings › Crawlers › “Block known artificial intelligence crawlers”.
What it does
Adds known AI crawlers to your robots.txt. It's off by default, and it doesn't remove anything already collected.

Squarespace itself warns that blocking might reduce your site traffic. Its list includes an outdated Anthropic name and leaves out some newer bots, so look at your robots.txt afterwards.

Where we'd start

A sensible starting policy

For most businesses that want customers to find and book them. Adjust it to your content and your server.

VisitorOur defaultWhy
Search enginesLet inGooglebot, Bingbot and Applebot bring customers, and feed Google's and Copilot's AI answers.
AI search botsLet inThey put you in ChatGPT, Claude and Perplexity answers, with a link.
User requests and signed agentsLet in, with a rate limitA person is behind each one. A limit protects you from a runaway agent.
Training crawlersYour callBlock if you'd rather your content didn't train AI; it costs little traffic. Allow if you want to be part of what models learn.
Bots that hide who they areRate-limit or blockAt your host or network. robots.txt won't reach them.
Read your robots.txt

Open yoursite.com/robots.txt. It's what you're asking bots today, whether you wrote it or your platform did.

Check your host's AI settings

Especially on Cloudflare: make sure blocking training hasn't blocked Googlebot or Bingbot.

Make the training decision once

Use the bouncer above, or your platform's switch, and write the date down.

Add a rate limit

Set well above what a busy customer does. It's your protection against the floods robots.txt can't stop.

Look at your logs monthly

Paste unfamiliar visitors into the directory above, and check their addresses before trusting a name.

Still open

What's still changing

The rules for who may visit your website are being written right now. These are the ones to watch.

DraftWeb Bot Auth

The standard behind signed agents was adopted by an IETF working group on 1 September 2026. It isn't final, and Anthropic, Perplexity, Meta and Microsoft don't sign yet.

DraftA shared vocabulary for AI preferences

The IETF is writing standard labels such as train-ai and search, so one line can express your wishes to every bot. Also still a draft.

Closed betaBeing paid for crawling

Cloudflare's pay per crawl, which charges AI crawlers per page, has been in closed beta since July 2025.

In courtAgents in court

Amazon v. Perplexity is back in the district court. The appeals court said agents with more autonomy could be treated differently.

EarlyAgents that pay

Visa's Trusted Agent Protocol and Mastercard's Agent Pay build on signed agents, so shops can tell an agent browsing from one paying.

OpenAgents in people's own browsers

Gemini in Chrome, Claude in Chrome and Copilot in Edge have no documented way to identify themselves to websites.

The bigger picture

Your front door now has two kinds of guest

For twenty years, the question was simple: let search engines in, keep the rest out. Now the machines at your door include your customers' own assistants, and turning them away looks a lot like turning the customer away.

The businesses that get this right won't block “AI”. They'll decide what their content is for, let in the visitors that bring customers, and protect their server from the ones that only take. We'll update this guide as the names and standards change, and the date at the top will tell you when we last checked.

DL
Written byDirk LaudonFounder, SproutMedia

AI specialist with a master's degree in Business Informatics, studied in Germany and Sweden with a focus on artificial intelligence. Dirk has spent years building software and apps, and has worked hands-on in e-commerce, affiliate marketing and social media, so he knows how customers find and choose a business online. At SproutMedia he helps businesses get ready for their customers' AI assistants and agents.

LinkedIn

SproutMedia writes about AI agents and agent readiness for businesses in Ireland, Europe, the United States and beyond.

Sources

  1. OpenAI: overview of OpenAI crawlers
  2. OpenAI: allowlisting ChatGPT's cloud browser
  3. Anthropic: does Anthropic crawl data from the web, and how can site owners block the crawler?
  4. Anthropic: using Claude in Chrome safely
  5. Google: common crawlers (Googlebot, Google-Extended)
  6. Google: Google-Agent
  7. Google: user-triggered fetchers
  8. Google Search Console: generative AI control
  9. Google Chrome Help: auto browse with Gemini
  10. Microsoft Bing: robots meta tags for Bing Chat and Copilot (September 2023)
  11. Perplexity: Perplexity crawlers
  12. Meta: web crawlers
  13. Apple: about Applebot
  14. Amazon: Amazonbot
  15. Common Crawl: CCBot
  16. RFC 9309: Robots Exclusion Protocol
  17. IETF: Web Bot Auth working group
  18. IETF: AI Preferences working group
  19. Cloudflare: Content Independence Day (1 July 2025)
  20. Cloudflare: signed agents (28 August 2025)
  21. Cloudflare: AI crawler traffic by purpose and industry (August 2025)
  22. Cloudflare: Content Independence Day, one year on (1 July 2026)
  23. Cloudflare: Radar 2025 Year in Review
  24. Cloudflare: Perplexity is using stealth, undeclared crawlers (4 August 2025)
  25. Perplexity: agents or bots? (4 August 2025)
  26. Cloudflare Community: Googlebot blocked by “Block AI training crawlers” (August 2026)
  27. Imperva: 2026 Bad Bot Report
  28. Digital Commerce 360: Adobe data on AI referral traffic, July 2026 (19 August 2026)
  29. Read the Docs: AI crawlers need to be more respectful (25 July 2024)
  30. TechCrunch: how OpenAI's bot crushed a seven-person company's website (10 January 2025)
  31. SourceHut: please stop externalizing your costs directly into my face (17 March 2025)
  32. TechCrunch: Meta's AI agent has been blocked from using Amazon.com (21 September 2026)
  33. Squarespace Help: blocking AI crawlers
  34. Shopify Help: editing robots.txt.liquid

Checked on 29 September 2026. Traffic figures come from different companies using different methods, and aren't directly comparable. Crawler names, rules and platform settings change often; check the companies' own pages before you rely on them. General information, not legal or security advice.