AI Crawler Registry
21 documented tokens. Registry reviewed July 26, 2026.
Search engines and AI products use different crawler names for different jobs. This source-backed registry separates documented search, user-requested retrieval, advertising review, model training, AI-use controls, and public-dataset collection.
Security boundary: This is a read-only reference built from public vendor documentation. Viewing it does not scan a website, change robots.txt, alter a firewall, or grant any crawler access. A crawler name can be spoofed, so verify vendor IP ranges or DNS before creating firewall rules. Robots.txt is not access control for private content.
Download the registry as JSON
Documented tokens by purpose
“Allowed” is not a universal recommendation. Search inclusion, user-requested retrieval, and model training are separate choices. Review the vendor source and your existing restricted paths before changing a policy.
Search
User retrieval
| Robots token | Vendor | Documented purpose | Source check |
Claude-User |
Anthropic |
Claude user-requested retrieval |
Vendor documentation Verified 2026-07-23 |
MistralAI-User |
Mistral |
Mistral Vibe user-requested page visits |
Vendor documentation Verified 2026-07-23 |
Amzn-User |
Amazon |
Amazon and Alexa user-requested retrieval |
Vendor documentation Verified 2026-07-26 |
Ads verification
| Robots token | Vendor | Documented purpose | Source check |
OAI-AdsBot |
OpenAI |
ChatGPT Ads landing-page validation |
Vendor documentation Verified 2026-07-23 |
AI training
AI use control
| Robots token | Vendor | Documented purpose | Source check |
Google-Extended |
Google |
Gemini training and grounding outside Search |
Vendor documentation Verified 2026-07-23 |
Applebot-Extended |
Apple |
Training use of content collected by Applebot |
Vendor documentation Verified 2026-07-23 |
Public dataset
| Robots token | Vendor | Documented purpose | Source check |
CCBot |
Common Crawl |
Common Crawl public web corpus |
Vendor documentation Verified 2026-07-23 |
Why some names are not assigned a result
DomainOptic does not convert product names, unverified bot lists, or ambiguous user-triggered behavior into guessed robots.txt results.
- ChatGPT-User: OpenAI's current publisher guidance does not document a dependable robots.txt allow-or-block interpretation for user-initiated requests.
- Perplexity-User: Perplexity's current first-party pages do not clearly reconcile robots.txt behavior for user-requested visits, so DomainOptic does not assign a badge.
- Kimi-User: Kimi states that robots.txt may not directly apply to user-triggered visits, so DomainOptic does not assign a badge.
- Qwen, DeepSeek, and Grok products: DomainOptic has not found first-party documentation for a distinct robots.txt token and purpose that can be reported without guessing.
Check your published policy
The Search and AI Visibility check reads a public homepage and robots.txt file. It does not impersonate a vendor crawler or change the scanned site.
Open Search and AI Visibility