AI crawler directory
Every AI crawler, and what it actually does.
The 50 AI bots TrustTraffic detects — the ones fetching your pages to answer people, indexing you for answer engines, and crawling you for model training. Search to check a user-agent, or filter by what the bot is for.
Indexes pages for ChatGPT search results and source citations.
Fetches a page live when a ChatGPT user opens or references it.
Crawls the open web to train OpenAI's models.
Indexes pages so Claude can cite them in answers.
Fetches a page live when a Claude user asks about it.
Crawls the web to train Anthropic's Claude models.
Fetches a page live to answer a Perplexity user's question.
Indexes pages for Perplexity's answer engine and citations.
Controls whether your content trains Gemini and Vertex AI.
Fetches pages for Vertex AI agents at a customer request.
Fetches sources a user adds to NotebookLM.
General-purpose Google crawler for research and product R&D.
Controls whether Applebot data trains Apple's models.
Indexes pages for Siri and Spotlight suggestions.
Bing's crawler; its index powers Copilot answers.
Crawls the web to train Meta's Llama models.
Meta crawler used to improve language models and products.
Crawls aggressively to train ByteDance / TikTok models.
Collects web data for TikTok / ByteDance AI systems.
Amazon crawler used to improve Alexa and search answers.
Fetches pages on demand for Amazon Bedrock AI agents.
Builds the open Common Crawl dataset widely used to train LLMs.
Collects web data to train Cohere's models.
Indexes pages for the You.com answer engine.
Fetches pages for DuckDuckGo's AI-assisted answers.
Fetches a page live for Le Chat (Mistral) users.
Extracts structured data from pages for knowledge graphs.
Collects web data for AI2's open research datasets.
Crawls images for ImageSift's visual search dataset.
Collects web data resold for AI training.
Webz.io crawler feeding AI-training datasets.
Crawls for Timpi's decentralized search index.
Huawei crawler for Petal Search and AI features.
Collects web data to train Huawei's PanGu models.
Crawls for Kagi's independent search index.
Collects web data resold for LLM training.
Indexes pages for the iAsk answer engine.
On-demand fetcher used by AI apps built on Firecrawl.
Semrush crawler collecting content for AI features.
Collects Japanese-language web data for LLM training.
Crawls and archives pages, resold for AI use.
Collects web data for AI research.
Gathers company data for aiHit's datasets.
Fetches pages for the Andi answer engine.
Fetches pages for Qualified's sales AI.
Yandex crawler for supplementary / AI data.
Collects web pages for machine-learning datasets.
Bulk-downloads images to build AI-training datasets.
Collects Japanese web data for LLM training.
Managed crawler used by AI apps to fetch pages.
Which of these are hitting your site?
The directory tells you what each bot is. TrustTraffic tells you which ones actually crawl your pages, how often, and where.