All crawlers
CCBot
Operated by Common Crawl · Training crawlers
Builds the open web-crawl datasets many models train on.
User-agent token
CCBotrobots.txt
Documented to honour robots.txt
Allow or block CCBot
By default every well-behaved crawler is allowed. To block CCBot specifically, add this to your robots.txt:
User-agent: CCBot Disallow: /
Official Common Crawl documentation
Did CCBot read your content?
Quillly detects CCBot and every other crawler server-side when they fetch your published pages, and shows it in your analytics — so you can see exactly which AI picked up which post, and when.
Track your crawlers free