AI crawlers, search bots & training agents
A reference of the 52+ crawler identities that visit your content — what each one is for, how it identifies itself, and whether it respects robots.txt. Quillly detects these on your published pages so you can see exactly which AI actually read you.
AI answers
User-triggered fetches when someone asks an AI assistant about a page. These are the visits that turn into citations — real readers sent your way.
Search indexes
Discovery and indexing crawlers that keep AI search and traditional search results fresh and eligible to surface your content.
Training crawlers
Collect public content to train and improve AI models. Blocking these is a robots.txt choice; being included shapes what models "know".
Other AI bots
Link previews, ads, and miscellaneous AI-operated fetchers you may see in your logs.
Fetches a page when ChatGPT answers a user’s question and cites it.
ChatGPT-UserAI answersFetches a page when Claude answers a user’s question and cites it.
Claude-UserAI answersReal-time fetch for answers and citations in Perplexity.
Perplexity-UserAI answersSupports the NotebookLM answer workflow.
Google-NotebookLMAI answersRead-aloud and assistant experiences.
Google-Read-AloudAI answersUser-triggered fetch for Le Chat answers.
MistralAI-UserAI answersMicrosoft Copilot AI answer fetches.
CopilotAI answersFresh answers in Alexa and Amazon products.
Amzn-UserAI answersReal-time AI-assisted answers with citations.
DuckAssistBotAI answersUser-triggered fetcher for Grok answers.
xAI-SearchBotAI answersDeep-research fetches for Grok answers.
Grok-DeepSearchAI answersFetches when users request or share URLs in Meta AI.
meta-externalfetcherAI answersUser-triggered fetcher for Kimi answers.
Kimi-UserAI answersUser-triggered fetcher for Qwen answers.
Qwen-UserAI answersSurfaces your site in ChatGPT search results.
OAI-SearchBotSearch indexContent discovery for Claude search.
Claude-SearchBotSearch indexMaintains the freshness of Perplexity’s answer index.
PerplexityBotSearch indexSearch Console URL inspection and answer-index crawling.
Google-InspectionToolSearch indexGoogle search indexing and discovery; feeds AI Overviews.
GooglebotSearch indexSearch / answer-index crawler for Mistral.
MistralAI-IndexSearch indexBing search indexing — also feeds Copilot and IndexNow engines.
BingbotSearch indexMakes content eligible for Amazon search surfaces.
Amzn-SearchBotSearch indexImproves Meta AI search result quality.
meta-webindexerSearch indexSearch / answer-index crawler for Kimi.
Kimi-SearchBotSearch indexSearch / answer-index crawler for TikTok.
TikTokSpiderSearch indexBaidu search indexing (China).
BaiduspiderSearch indexSearch / answer-index crawler for You.com.
YouBotSearch indexCollects public content to train future OpenAI models.
GPTBotTrainingCollects public content to improve Claude.
ClaudeBotTrainingGeneric public-content fetches for research and development.
GoogleOtherTrainingrobots.txt token controlling Gemini model training.
Google-ExtendedTrainingrobots.txt token controlling Apple AI training.
Applebot-ExtendedTrainingPowers Siri, Spotlight and Apple search + training.
ApplebotTrainingImproves Amazon products and services (incl. Alexa).
AmazonbotTrainingIndexes and improves Meta products and AI.
meta-externalagentTrainingPublic-content crawler for Kimi.
KimiBotTrainingPublic-content crawler for ByteDance models.
BytespiderTrainingPublic-content crawler for Baidu ERNIE.
ERNIEBotTrainingPublic-content crawler for Qwen.
QwenBotTrainingPublic-content crawler for ChatGLM.
ChatGLM-SpiderTrainingPublic-content crawler for DeepSeek.
DeepSeekBotTrainingPublic-content crawler for Cohere.
cohere-aiTrainingFinds documents for AI research systems.
AI2BotTrainingBuilds the open web-crawl datasets many models train on.
CCBotTrainingAd and landing-page fetch workflows.
OAI-AdsBotOtherGeneral xAI Grok fetcher.
GrokBotOtherGeneral xAI web crawler.
xAI-Web-CrawlerOtherAdvertising and business-product fetches.
meta-externaladsOtherGenerates shared-link previews on Facebook.
facebookexternalhitOtherFetcher for ByteDance’s Doubao assistant.
DoubaobotOtherFetcher for Baidu’s Ernie Bot (Yiyan).
YiyanBotOtherFetcher for Alibaba’s Tongyi assistant.
TongyiBotOtherWhat is an AI crawler?
An AI crawler is an automated agent that requests web pages on behalf of an AI company. Some collect public content to train models (GPTBot, ClaudeBot, Google-Extended), some keep an answer index fresh so your pages can be cited (OAI-SearchBot, PerplexityBot, Claude-SearchBot), and some fetch a page in real time when a user asks an assistant a question (ChatGPT-User, Claude-User, Perplexity-User). Each identifies itself with a distinct user-agent string.
Why AI crawler visits matter
Getting crawled by these agents is how your content becomes eligible to be cited in ChatGPT, Claude, Perplexity, Google AI Overviews and the rest. A real-time answer bot (like ChatGPT-User) fetching your page usually means an actual person is about to read a summary of it with a link back to you. Knowing which bots visit — and which pages they read — tells you whether your content is reaching the AI answer layer at all.
How Quillly uses this directory
Standard analytics run in the browser, so they never see crawlers — bots don't execute JavaScript. Quillly detects these agents server-side when they request your published pages, then shows the crawls in your analytics: which AI read which page, and when. It's the difference between hoping you're in the training set and seeing GPTBot fetch your post an hour after you publish it.
Allowing or blocking crawlers
You control access with robots.txt. To allow every well-behaved crawler, do nothing. To block a specific one, add a rule for its user-agent token — for example User-agent: GPTBot then Disallow: /. Crawlers marked with a shield in the directory are documented to honour robots.txt.
See which AI actually read your content
Quillly publishes to your domain and tracks every crawler that visits — so you know the moment ChatGPT, Claude or Perplexity picks up a post.