Early-adopter pricing — $9/mo, before it rises.See plan →
Crawler directory

AI crawlers, search bots & training agents

A reference of the 52+ crawler identities that visit your content — what each one is for, how it identifies itself, and whether it respects robots.txt. Quillly detects these on your published pages so you can see exactly which AI actually read you.

AI answers

User-triggered fetches when someone asks an AI assistant about a page. These are the visits that turn into citations — real readers sent your way.

Search indexes

Discovery and indexing crawlers that keep AI search and traditional search results fresh and eligible to surface your content.

Training crawlers

Collect public content to train and improve AI models. Blocking these is a robots.txt choice; being included shapes what models "know".

Other AI bots

Link previews, ads, and miscellaneous AI-operated fetchers you may see in your logs.

ChatGPT-User
OpenAI

Fetches a page when ChatGPT answers a user’s question and cites it.

ChatGPT-UserAI answers
Claude-User
Anthropic

Fetches a page when Claude answers a user’s question and cites it.

Claude-UserAI answers
Perplexity-User
Perplexity

Real-time fetch for answers and citations in Perplexity.

Perplexity-UserAI answers
Google-NotebookLM
Google

Supports the NotebookLM answer workflow.

Google-NotebookLMAI answers
Google-Read-Aloud
Google

Read-aloud and assistant experiences.

Google-Read-AloudAI answers
MistralAI-User
Mistral

User-triggered fetch for Le Chat answers.

MistralAI-UserAI answers
Copilot
Microsoft

Microsoft Copilot AI answer fetches.

CopilotAI answers
Amzn-User
Amazon

Fresh answers in Alexa and Amazon products.

Amzn-UserAI answers
DuckAssistBot
DuckDuckGo

Real-time AI-assisted answers with citations.

DuckAssistBotAI answers
xAI-SearchBot
xAI

User-triggered fetcher for Grok answers.

xAI-SearchBotAI answers
Grok-DeepSearch
xAI

Deep-research fetches for Grok answers.

Grok-DeepSearchAI answers
meta-externalfetcher
Meta

Fetches when users request or share URLs in Meta AI.

meta-externalfetcherAI answers
Kimi-User
Moonshot AI

User-triggered fetcher for Kimi answers.

Kimi-UserAI answers
Qwen-User
Alibaba

User-triggered fetcher for Qwen answers.

Qwen-UserAI answers
OAI-SearchBot
OpenAI

Surfaces your site in ChatGPT search results.

OAI-SearchBotSearch index
Claude-SearchBot
Anthropic

Content discovery for Claude search.

Claude-SearchBotSearch index
PerplexityBot
Perplexity

Maintains the freshness of Perplexity’s answer index.

PerplexityBotSearch index
Google-InspectionTool
Google

Search Console URL inspection and answer-index crawling.

Google-InspectionToolSearch index
Googlebot
Google

Google search indexing and discovery; feeds AI Overviews.

GooglebotSearch index
MistralAI-Index
Mistral

Search / answer-index crawler for Mistral.

MistralAI-IndexSearch index
Bingbot
Microsoft

Bing search indexing — also feeds Copilot and IndexNow engines.

BingbotSearch index
Amzn-SearchBot
Amazon

Makes content eligible for Amazon search surfaces.

Amzn-SearchBotSearch index
meta-webindexer
Meta

Improves Meta AI search result quality.

meta-webindexerSearch index
Kimi-SearchBot
Moonshot AI

Search / answer-index crawler for Kimi.

Kimi-SearchBotSearch index
TikTokSpider
ByteDance

Search / answer-index crawler for TikTok.

TikTokSpiderSearch index
Baiduspider
Baidu

Baidu search indexing (China).

BaiduspiderSearch index
YouBot
You.com

Search / answer-index crawler for You.com.

YouBotSearch index
GPTBot
OpenAI

Collects public content to train future OpenAI models.

GPTBotTraining
ClaudeBot
Anthropic

Collects public content to improve Claude.

ClaudeBotTraining
GoogleOther
Google

Generic public-content fetches for research and development.

GoogleOtherTraining
Google-Extended
Google

robots.txt token controlling Gemini model training.

Google-ExtendedTraining
Applebot-Extended
Apple

robots.txt token controlling Apple AI training.

Applebot-ExtendedTraining
Applebot
Apple

Powers Siri, Spotlight and Apple search + training.

ApplebotTraining
Amazonbot
Amazon

Improves Amazon products and services (incl. Alexa).

AmazonbotTraining
meta-externalagent
Meta

Indexes and improves Meta products and AI.

meta-externalagentTraining
KimiBot
Moonshot AI

Public-content crawler for Kimi.

KimiBotTraining
Bytespider
ByteDance

Public-content crawler for ByteDance models.

BytespiderTraining
ERNIEBot
Baidu

Public-content crawler for Baidu ERNIE.

ERNIEBotTraining
QwenBot
Alibaba

Public-content crawler for Qwen.

QwenBotTraining
ChatGLM-Spider
Zhipu AI

Public-content crawler for ChatGLM.

ChatGLM-SpiderTraining
DeepSeekBot
DeepSeek

Public-content crawler for DeepSeek.

DeepSeekBotTraining
cohere-ai
Cohere

Public-content crawler for Cohere.

cohere-aiTraining
AI2Bot
Allen AI

Finds documents for AI research systems.

AI2BotTraining
CCBot
Common Crawl

Builds the open web-crawl datasets many models train on.

CCBotTraining
OAI-AdsBot
OpenAI

Ad and landing-page fetch workflows.

OAI-AdsBotOther
GrokBot
xAI

General xAI Grok fetcher.

GrokBotOther
xAI-Web-Crawler
xAI

General xAI web crawler.

xAI-Web-CrawlerOther
meta-externalads
Meta

Advertising and business-product fetches.

meta-externaladsOther
facebookexternalhit
Meta

Generates shared-link previews on Facebook.

facebookexternalhitOther
Doubaobot
ByteDance

Fetcher for ByteDance’s Doubao assistant.

DoubaobotOther
YiyanBot
Baidu

Fetcher for Baidu’s Ernie Bot (Yiyan).

YiyanBotOther
TongyiBot
Alibaba

Fetcher for Alibaba’s Tongyi assistant.

TongyiBotOther

What is an AI crawler?

An AI crawler is an automated agent that requests web pages on behalf of an AI company. Some collect public content to train models (GPTBot, ClaudeBot, Google-Extended), some keep an answer index fresh so your pages can be cited (OAI-SearchBot, PerplexityBot, Claude-SearchBot), and some fetch a page in real time when a user asks an assistant a question (ChatGPT-User, Claude-User, Perplexity-User). Each identifies itself with a distinct user-agent string.

Why AI crawler visits matter

Getting crawled by these agents is how your content becomes eligible to be cited in ChatGPT, Claude, Perplexity, Google AI Overviews and the rest. A real-time answer bot (like ChatGPT-User) fetching your page usually means an actual person is about to read a summary of it with a link back to you. Knowing which bots visit — and which pages they read — tells you whether your content is reaching the AI answer layer at all.

How Quillly uses this directory

Standard analytics run in the browser, so they never see crawlers — bots don't execute JavaScript. Quillly detects these agents server-side when they request your published pages, then shows the crawls in your analytics: which AI read which page, and when. It's the difference between hoping you're in the training set and seeing GPTBot fetch your post an hour after you publish it.

Allowing or blocking crawlers

You control access with robots.txt. To allow every well-behaved crawler, do nothing. To block a specific one, add a rule for its user-agent token — for example User-agent: GPTBot then Disallow: /. Crawlers marked with a shield in the directory are documented to honour robots.txt.

See which AI actually read your content

Quillly publishes to your domain and tracks every crawler that visits — so you know the moment ChatGPT, Claude or Perplexity picks up a post.