An AI crawler is a program that an AI company sends out to read websites. Each one does one of three jobs. Some collect pages to train future models. Some build the index an AI search tool answers from. Others open a page because a person asked a question right now.
The job decides what blocking the bot costs you. Block GPTBot, OpenAI's training crawler, and ChatGPT search carries on as before. Block OAI-SearchBot, and ChatGPT stops showing your pages in its search answers.
Our advice is to leave every AI bot open, training crawlers included. What a model learns in training shapes how it describes your brand when it answers without searching. An assistant answering from memory can quote a price plan you dropped last year, the case our LLM SEO explainer walks through. A blocked bot is the cheapest AI visibility problem to fix, and it is the first thing we check.
What does each kind of AI crawler feed?
Training crawlers feed the next version of a model. They collect pages that may be used for training, and GPTBot and ClaudeBot are the best known.
Search crawlers feed the index, the store of pages an AI search tool pulls from while it writes an answer. OAI-SearchBot does this for ChatGPT, PerplexityBot for Perplexity and Claude-SearchBot for Claude.
User-triggered fetchers feed the answer being written right now. They open one page because a person asked for it, as ChatGPT-User, Claude-User and Perplexity-User do.
Letting these crawlers in is the first step of generative engine optimization (GEO), the work of getting a brand named and cited in AI answers.

What each kind of AI bot feeds, with the best-known names in each group.
Two classic search crawlers feed AI answers as well. Googlebot collects the pages behind AI Overviews and AI Mode, which Google treats as part of Search. Bingbot builds the Bing index, which ChatGPT still draws on beside its own.
The AI crawlers list: bots documented by the companies that run them
The AI crawlers behind ChatGPT, Claude and Perplexity are OpenAI's GPTBot, OAI-SearchBot and ChatGPT-User, Anthropic's ClaudeBot, Claude-SearchBot and Claude-User, and Perplexity's PerplexityBot and Perplexity-User. Google, Apple, Meta, Amazon and Microsoft document their bots too. Every bot in the table below has a page published by the company that runs it. Each row gives that company's own account of the bot's job and of what blocking it changes.
robots.txt is the small text file at the root of a site that tells bots which pages they may read. In that file, you write each bot's name as the first column gives it, and the robots.txt standard says capital letters do not matter.
Bot | Run by | Job | What it does | What you lose if you block it |
|---|---|---|---|---|
GPTBot | Training | Collects pages that may train OpenAI's models | A place in OpenAI's future training data; ChatGPT search is unaffected | |
OAI-SearchBot | OpenAI | Search | Finds pages to show in ChatGPT's search answers | Your pages in ChatGPT search answers |
ChatGPT-User | OpenAI | Fetches for a user | Opens a page when a ChatGPT user's request needs it | ChatGPT can no longer open your page while it answers someone, though the block may not be honored |
OAI-AdsBot | OpenAI | Fetches for a user | Checks the safety of pages submitted as ads on ChatGPT | Only matters if you advertise on ChatGPT, since it visits only pages submitted as ads |
ClaudeBot | Training | Collects pages that could help train Anthropic's models | Your future pages in Anthropic's training data | |
Claude-SearchBot | Anthropic | Search | Reads pages to improve Claude's search results | Your pages in Claude's search results; Anthropic says blocking may reduce visibility |
Claude-User | Anthropic | Fetches for a user | Opens a page when a Claude user asks a question | Claude stops retrieving your page for people's questions |
Googlebot | Search | Crawls for Google Search | Google Search itself, AI Overviews and AI Mode included | |
Google-Extended | Training and answers | A setting that Google's regular crawlers read; it controls whether Gemini may use your pages | Your pages in Gemini's training and in the Google Search results Gemini uses while it answers; Google Search is unaffected | |
Google-CloudVertexBot | Search | Crawls a site at its owner's request, for an agent built on Vertex AI, Google's cloud service for building AI agents | Only the crawls a site owner requests for their own Vertex AI agent; Google Search and other products are unaffected | |
Google-Agent | Fetches for a user | Lets agents on Google's systems browse and act for a user | Google's agents cannot browse or act on your site for a person, though the block may not be honored | |
Google-GeminiNotebook | Fetches for a user | Fetches pages a Gemini Notebook user adds as sources | Gemini Notebook users cannot add your page as a source, though the block may not be honored | |
PerplexityBot | Search | Surfaces and links websites in Perplexity search; Perplexity says it is not used for model training | Your site surfaced and linked in Perplexity answers | |
Perplexity-User | Perplexity | Fetches for a user | Visits a page when a Perplexity user's question needs it | Perplexity can no longer fetch your page when a person asks about it, though the block may not be honored |
Applebot | Search | Crawls for search in Siri, Spotlight and Safari; Apple says the data may also train its models | Your place in Siri, Spotlight and Safari search | |
Applebot-Extended | Apple | Training | A setting that crawls nothing; it controls whether Apple may train its models on your pages | Your pages in Apple's model training; Apple search results are unaffected |
Meta-ExternalAgent | Training | Crawls for uses such as training Meta's AI models or indexing content for its products | Your pages in Meta's AI training and in Meta's product indexing | |
Meta-WebIndexer | Meta | Search | Reads pages to improve Meta AI's search results | Citations and links to your pages in Meta AI answers |
Meta-ExternalFetcher | Meta | Fetches for a user | Fetches a link at a user's request | Meta cannot open a link to your site when someone asks it to, though the block may not be honored |
Amazonbot | Training | Improves Amazon's products and services; Amazon says the data may train its AI models | Your pages in Amazon's products and services, and possibly in its AI training | |
Amzn-SearchBot | Amazon | Search | Crawls for search in Amazon products such as Alexa; Amazon says it is not used for model training | Your place in Amazon search, such as Alexa answers |
Amzn-User | Amazon | Fetches for a user | Fetches live information for a user, such as for an Alexa question | Alexa cannot pull current facts from your page, though the block may not be honored |
Bingbot | Search | Crawls for Bing search | Bing search, and one of the outside sources ChatGPT still uses next to its own index |
OpenAI, Anthropic and Apple let you close training and keep every bot that puts you in answers. At Google, Amazon and Meta, the robots.txt name that covers training also covers Gemini's answers or the company's products.
Which other AI bots show up in server logs?
Peec AI, the tracking tool we use, recognizes 26 more AI bots in server logs, and the largest share of them collects training data. None of them needs a line of its own in robots.txt. The catch-all User-agent: * group already lets them in, and that is where we leave them.
Job, as our tracking tool classifies it | Bots (operator in parentheses where it documents the bot) |
|---|---|
Training: collecting pages for AI models | CCBot (Common Crawl), AI2Bot (Allen Institute for AI), Brightbot (Bright Data), Bytespider, DeepSeekBot, cohere-ai, PanguBot, GrokBot, Timpibot, Webzio-Extended, Diffbot, FacebookBot |
Search: building an index for AI answers | AzureAI-SearchBot, xAI-Grok, Grok-DeepSearch |
User-triggered: fetching a page for a person's question | MistralAI-User (Mistral AI), DuckAssistBot (DuckDuckGo), Gemini-Deep-Research, GoogleAgent-Mariner, NovaAct, Manus-User, Claude-Code |
Other or unclear | Claude-Web, YouBot, omgilibot, MyCentralAIScraperBot |
Two of the 26 are answer bots that link to their sources, according to the companies that run them. DuckDuckGo's DuckAssistBot fetches pages in real time for DuckDuckGo's AI-assisted answers. Mistral AI documents three bots: MistralAI-Training for model training, MistralAI-Index for its search and MistralAI-User, which may visit a page for a user's question and link to it.
What does a month of AI bot traffic look like?
On one B2B site we track, 17 AI bots made 1,391 verified visits in the month to 27 September 2026. Training crawlers made almost three in four of them. We counted each bot under the job its vendor documents, or under our tracking tool's label where the vendor gives no single job. Visits from bots that fake a name are left out. Split by job, the month looked like this:
Job | AI bots seen on the site | Verified visits in the month |
|---|---|---|
Training | 9 bots, led by ClaudeBot | 1,022 |
Search | 3 bots, led by xAI-Grok | 133 |
Fetches for a user | 3 bots, led by MistralAI-User | 88 |
Other or unclear | YouBot, and Google-CloudVertexBot, which crawls for site owners' own AI agents | 148 |

One B2B site we track, verified visits, 30 days to 27 September 2026.
ClaudeBot, Anthropic's training crawler, was the busiest visitor by far, with 539 visits, close to two in five.
Blocking the busiest bots on this site would have left ChatGPT's search answers as they were. Blocking OAI-SearchBot alone would have taken the site out of them. OAI-SearchBot and ChatGPT-User are the two bots that bring a page into ChatGPT's answers, one through search and one through a user's request. Between them they made 50 visits in the month. GPTBot, OpenAI's training crawler, made 12.
What robots.txt should you use for AI crawlers?
Allow every AI bot in robots.txt, and name the eight that bring pages into answers on ChatGPT, Claude, Perplexity, Meta AI and Alexa: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, meta-webindexer and Amzn-SearchBot. The file below does that, gives the training bots a group of their own and ends with a Sitemap line.
The search and user-triggered group is what lets AI assistants find and read you, so keep it open on every site. The training group is a business decision, and our advice is to leave it open too. Three names in it also cover answers or products: Google-Extended, Amazonbot and meta-externalagent. Closing Amazonbot or meta-externalagent also takes your pages out of Amazon's and Meta's product features. To close the group anyway, change its Allow: / line to Disallow: /.
Google-Extended controls whether Google may use your pages to train Gemini and to hand them to Gemini while it answers. It has no effect on Google Search, including AI Overviews and AI Mode. Google's regular crawlers read the name in robots.txt, so Google-Extended never visits your site. Applebot-Extended works the same way for Apple's model training.
The Disallow: /admin/ lines stand in for your own. They sit in every group because of one rule in the robots.txt standard, RFC 9309. A bot that finds its own name follows only that group and ignores the catch-all User-agent: * group. So copy each Disallow line you already have into every named group. Do the same with any rule you add later, since a line added only to the catch-all group reaches none of the bots you named.
On a site that blocks nothing today, the file's job is to name the AI bots. Once they are named, a Disallow: / added later to the catch-all group cannot lock them out. It also shows whoever edits the file next which bots were let in on purpose. Put your own domain in the Sitemap line.
llms.txt, a separate file, lists your main pages for AI tools and grants or blocks nothing, which is why we put it last among the technical fixes.
How do you check whether your CDN or firewall blocks AI bots?
To check whether Cloudflare is blocking AI crawlers, open AI Crawl Control in the Cloudflare dashboard, then read Security Events for AI bots that were challenged or blocked. Any other CDN has bot settings and a security log to read the same way. A CDN is the service in front of a site that speeds it up and filters its traffic. A robots.txt that welcomes every AI bot can still sit behind a firewall that turns them away.
On one client's site, Cloudflare's setting to block AI training bots was also turning away OpenAI's bots that fetch pages for people asking ChatGPT, and switching it off in the Cloudflare dashboard fixed it.
On Cloudflare, the AI defaults have changed more than once:
In July 2025, Cloudflare started asking every new domain at sign-up whether to allow AI crawlers, with blocking as the default.
In July 2026, Cloudflare announced that from 15 September 2026, new domains block training bots and agent bots on pages that show ads. Agent bots are the ones acting for a person in real time.
Bot Fight Mode challenges traffic that looks like a known bot, and Cloudflare says a custom firewall rule cannot skip it.
The riskiest of these settings is the training block, because on Cloudflare it can take a site out of search as well. Cloudflare says a site set to block training also blocks multi-purpose crawlers such as Googlebot, Applebot and Bingbot, the crawlers behind Google, Apple and Bing search. Leave training allowed in Cloudflare's security settings. If you want training bots out, close them in robots.txt instead, where each one has its own name and Googlebot stays open.
Run the checks in this order:
In your CDN, look for any setting that blocks AI bots or AI training, for Bot Fight Mode, and for firewall rules that match bot names. On Cloudflare, AI Crawl Control, available on all plans, shows which AI crawlers visit and lets you allow or block each one.
Read the security log, called Security Events on Cloudflare, for AI bots that were challenged or blocked.
Ask your host whether its own firewall or a security plugin filters bots, since those sit outside the CDN.
Run the same checks for Bingbot, and add your site to Bing Webmaster Tools, Microsoft's console for site owners, to see what Bing has indexed.
How can you see which AI bots visit your site?
Your server logs and your CDN's analytics record every bot visit by name, so they show which AI bots reach you and which get turned away. In a raw server log, search the user-agent field, the part of each line where the visitor names itself, for names such as OAI-SearchBot and PerplexityBot. The status code beside each visit gives the result: 200 means the bot got the page, and 403 or a challenge page means it was turned away.
Start with the bots behind ChatGPT, Perplexity and Claude answers, such as OAI-SearchBot, ChatGPT-User, PerplexityBot and Claude-SearchBot. A 403 next to any of them is the first thing to fix.
Questions people ask about AI crawlers
What is ClaudeBot?
ClaudeBot is Anthropic's web crawler: it collects public pages that could help train Anthropic's Claude models. Blocking it in robots.txt tells Anthropic to keep a site's future pages out of that training data. Claude's search and its live fetches for users run on Claude-SearchBot and Claude-User, which a ClaudeBot rule does not reach.
Should I block GPTBot?
For most businesses, no. GPTBot only collects pages that may train OpenAI's future models, and OpenAI runs it apart from OAI-SearchBot, so ChatGPT search is unaffected. Letting GPTBot in means OpenAI's next models can learn from your pages. What they learn feeds how ChatGPT talks about your brand when it answers from memory. Publishers who sell their writing are the exception, since their content is the product.
Does blocking AI bots protect my content?
Only partly. robots.txt works from now on, and bots follow it by choice. The big crawlers do: OpenAI treats a GPTBot block as a request to keep your content out of training, and Anthropic, Google, Apple and Common Crawl say their crawlers follow the file. Bots that fetch a page for a person may skip it, and copies of your text on other websites sit outside it.
What is OAI-SearchBot?
OAI-SearchBot is OpenAI's search crawler: it reads pages so ChatGPT can show them in its search answers. OpenAI says sites that opt out of it "will not be shown in ChatGPT search answers, though can still appear as navigational links." It runs apart from GPTBot, OpenAI's training crawler, so you can allow one and block the other. Allowing OAI-SearchBot is step one of getting cited in ChatGPT.
How do I verify a bot is real?
Check the visitor's IP address against the list its vendor publishes, because anyone can fake a bot's name. OpenAI, Anthropic, Google, Perplexity, Apple, Amazon and Microsoft publish one; Meta's crawler page gives none. Google and Apple support a reverse DNS check: the name behind the IP should belong to the vendor, such as googlebot.com or applebot.apple.com. A lookup of that name should then return the same IP.
Does robots.txt apply to ChatGPT-User?
Not reliably. OpenAI says ChatGPT-User acts on a user's request, so "robots.txt rules may not apply" to it. Perplexity says Perplexity-User generally ignores robots.txt, and Google, Meta and Amazon say much the same about their user-triggered fetchers. Anthropic, Mistral AI and DuckDuckGo say their bots follow it. For the fetchers that skip it, a hard block has to happen in your firewall, and we would not add one.
How often do AI bots crawl a site?
On one B2B site we track, ClaudeBot came about 18 times a day over a month, while ChatGPT-User came four times in the whole month. If ClaudeBot visits too often, Anthropic supports a Crawl-delay line in robots.txt to slow it down. OpenAI says ChatGPT search can take about 24 hours to adjust to a robots.txt change, and Perplexity says its systems may take up to 24 hours.
What does it cost to have AI bot access checked?
Checking AI bot access yourself costs nothing: read your robots.txt, your CDN's bot settings and your server logs. We check all three in the first month of every GEO engagement, together with the other technical fixes and the list of buyer questions we track. Agenzy's GEO service starts from 5,000 EUR a month, ex VAT, and the full scope is on our GEO service page.
About Agenzy
Agenzy is a GEO agency based in Vilnius, working with brands in the US, the UK and across Europe. We get brands named and recommended inside ChatGPT, Gemini, Google AI Overviews, Perplexity, Claude and Copilot. We are an official Peec AI partner. As of September 2026: 500,000+ AI chats analysed, 15,000+ prompts tracked, 150+ audits completed, 1,000,000+ EUR generated for clients by AI search. Dated cases sit at agenzy.lt/case-studies.




