AI crawler directory

Every AI crawler, and what it actually does.

The 50 AI bots TrustTraffic detects — the ones fetching your pages to answer people, indexing you for answer engines, and crawling you for model training. Search to check a user-agent, or filter by what the bot is for.

50 crawlers
OAI-SearchBot
OpenAI
indexing

Indexes pages for ChatGPT search results and source citations.

ChatGPT-User
OpenAI
answers

Fetches a page live when a ChatGPT user opens or references it.

GPTBot
OpenAI
training

Crawls the open web to train OpenAI's models.

Claude-SearchBot
Anthropic
indexing

Indexes pages so Claude can cite them in answers.

Claude-User
Anthropic
answers

Fetches a page live when a Claude user asks about it.

ClaudeBot
Anthropic
training

Crawls the web to train Anthropic's Claude models.

Perplexity-User
Perplexity
answers

Fetches a page live to answer a Perplexity user's question.

PerplexityBot
Perplexity
indexing

Indexes pages for Perplexity's answer engine and citations.

Google-Extended
Google
training

Controls whether your content trains Gemini and Vertex AI.

Google-CloudVertexBot
Google
training

Fetches pages for Vertex AI agents at a customer request.

Google-NotebookLM
Google
answers

Fetches sources a user adds to NotebookLM.

GoogleOther
Google
other

General-purpose Google crawler for research and product R&D.

Applebot-Extended
Apple
training

Controls whether Applebot data trains Apple's models.

Applebot
Apple
indexing

Indexes pages for Siri and Spotlight suggestions.

Bingbot
Microsoft
indexing

Bing's crawler; its index powers Copilot answers.

Meta-ExternalAgent
Meta
training

Crawls the web to train Meta's Llama models.

FacebookBot
Meta
other

Meta crawler used to improve language models and products.

Bytespider
ByteDance
training

Crawls aggressively to train ByteDance / TikTok models.

TikTokSpider
ByteDance
training

Collects web data for TikTok / ByteDance AI systems.

Amazonbot
Amazon
other

Amazon crawler used to improve Alexa and search answers.

bedrockbot
Amazon
answers

Fetches pages on demand for Amazon Bedrock AI agents.

CCBot
Common Crawl
training

Builds the open Common Crawl dataset widely used to train LLMs.

cohere-ai
Cohere
training

Collects web data to train Cohere's models.

YouBot
You.com
indexing

Indexes pages for the You.com answer engine.

DuckAssistBot
DuckDuckGo
indexing

Fetches pages for DuckDuckGo's AI-assisted answers.

MistralAI-User
Mistral
answers

Fetches a page live for Le Chat (Mistral) users.

Diffbot
Diffbot
other

Extracts structured data from pages for knowledge graphs.

AI2Bot
Allen Institute for AI
training

Collects web data for AI2's open research datasets.

ImagesiftBot
ImageSift
other

Crawls images for ImageSift's visual search dataset.

Omgilibot
Webz.io
training

Collects web data resold for AI training.

Webzio-Extended
Webz.io
training

Webz.io crawler feeding AI-training datasets.

Timpibot
Timpi
indexing

Crawls for Timpi's decentralized search index.

PetalBot
Huawei
indexing

Huawei crawler for Petal Search and AI features.

PanguBot
Huawei
training

Collects web data to train Huawei's PanGu models.

Kagibot
Kagi
indexing

Crawls for Kagi's independent search index.

VelenPublicWebCrawler
Velen
training

Collects web data resold for LLM training.

iaskspider
iAsk
indexing

Indexes pages for the iAsk answer engine.

FirecrawlAgent
Firecrawl
other

On-demand fetcher used by AI apps built on Firecrawl.

SemrushBot-OCOB
Semrush
training

Semrush crawler collecting content for AI features.

SBIntuitionsBot
SB Intuitions
training

Collects Japanese-language web data for LLM training.

Panscient
Panscient
other

Crawls and archives pages, resold for AI use.

PoseidonResearchCrawler
Poseidon Research
training

Collects web data for AI research.

aiHitBot
aiHit
other

Gathers company data for aiHit's datasets.

Andibot
Andi
indexing

Fetches pages for the Andi answer engine.

QualifiedBot
Qualified
other

Fetches pages for Qualified's sales AI.

YandexAdditional
Yandex
other

Yandex crawler for supplementary / AI data.

FriendlyCrawler
FriendlyCrawler
training

Collects web pages for machine-learning datasets.

img2dataset
img2dataset
training

Bulk-downloads images to build AI-training datasets.

Cotoyogi
ROIS
training

Collects Japanese web data for LLM training.

Crawlspace
Crawlspace
other

Managed crawler used by AI apps to fetch pages.

Which of these are hitting your site?

The directory tells you what each bot is. TrustTraffic tells you which ones actually crawl your pages, how often, and where.