AI Crawler List & Verification (2026)
Who to allow, who retrieves, who is user-triggered — and how to verify identity with official IP ranges and reverse DNS (user agents can be spoofed).
Start with the three categories
A vendor often runs several crawlers by purpose — blocking the training bot does not affect its search or user-triggered bots. Mixing them up is the most common cause of accidental blocking:
| Category | What it does | Example tokens | Recommendation |
|---|---|---|---|
| Training | Collects content for model training | GPTBot · ClaudeBot · Google-Extended · CCBot · Applebot-Extended · Bytespider · meta-externalagent | Your content policy (we allow: public content wants to be known) |
| Search / retrieval | Builds an index so the assistant can cite your pages | OAI-SearchBot · PerplexityBot · Claude-SearchBot · MistralAI-Index | Allow — this is the direct doorway to being cited |
| User-triggered | Fetches in real time for a user’s question | ChatGPT-User · Perplexity-User · Claude-User · MistralAI-User | Allow — a user is asking; if you are unreachable you lose the moment |
Collects content for model training
Your content policy (we allow: public content wants to be known)
Builds an index so the assistant can cite your pages
Allow — this is the direct doorway to being cited
Fetches in real time for a user’s question
Allow — a user is asking; if you are unreachable you lose the moment
The list (with official sources)
Every token below is tied to an official source; where none exists we say so plainly (observed only, not officially confirmed) instead of guessing.
| token / user agent | Operator | Category | Official source |
|---|---|---|---|
| GPTBot | OpenAI | Training | openai.com/gptbot.json (IP ranges) |
| OAI-SearchBot | OpenAI | Search | openai.com/searchbot.json |
| ChatGPT-User | OpenAI | User-triggered | openai.com/chatgpt-user.json |
| ClaudeBot | Anthropic | Training | claude.com/crawling/bots.json (IP ranges) |
| Claude-SearchBot | Anthropic | Search | same (three-bot framework documented) |
| Claude-User | Anthropic | User-triggered | same |
| PerplexityBot | Perplexity | Search | perplexity.ai/perplexitybot.json |
| Perplexity-User | Perplexity | User-triggered | perplexity.ai/perplexity-user.json |
| Googlebot | Search | developers.google.com/static/crawling/ipranges/common-crawlers.json | |
| Google-Extended | Training (token) | same — training only, does not affect Search ranking | |
| Google-Agent | User-triggered | same (added 2026-03) | |
| Bingbot | Microsoft | Search (also powers Copilot) | bing.com/toolbox/bingbot.json |
| Applebot / Applebot-Extended | Apple | Search / training token | search.developer.apple.com/applebot.json |
| Amazonbot / Amzn-SearchBot / Amzn-User | Amazon | Training / Search / User-triggered | developer.amazon.com/amazonbot (rDNS is Amazon’s recommended check) |
| CCBot | Common Crawl | Training (open corpus) | official convention: rDNS to *.commoncrawl.org |
| meta-externalagent / Meta-ExternalFetcher | Meta | Training / user-triggered | developers.facebook.com/docs/sharing/webmasters/crawler |
| Bytespider | ByteDance (Doubao / TikTok search) | Training | zhanzhang.toutiao.com/docs/intro/26899 (official; rDNS suffix .byted.org) |
| DoubaoBot | ByteDance (Doubao) | AI Agent | listed in industry directories; no dedicated official page found (distinct from Bytespider) |
China-based AI crawlers
Chinese vendors publish less than their international peers (often no IP-range JSON). Official sources are marked; where there is none, we say so.
| token / user agent | Operator | Category | Official source |
|---|---|---|---|
| Baiduspider / Baiduspider-render | Baidu | Search / rendering | www.baidu.com/robots.txt (official declaration) |
| Sogou web spider / Sogou inst spider | Sogou (Tencent) | Search | www.sogou.com/robots.txt (official declaration) |
| 360Spider | 360 | Search | www.so.com/help/spider_ip.html — six official IP ranges published, with a note asking sites not to block it; officially states it does not support nslookup (IP ranges only) |
| Kimi-SearchBot | Moonshot AI (Kimi) | Search | kimi.ai/policies/kimi-crawlers (official policy page) |
| DeepSeekBot | DeepSeek | Training | no official crawler page found; industry directories record it as not honoring robots.txt |
| ChatGLM-Spider | Zhipu | Training / scraping | no official page found (industry entry is also incomplete) |
| QwenBot / TongyiBot / Qwen-User | Alibaba (Qwen) | Training / Search / User-triggered | sparse official English docs; recorded in industry directories |
| PetalBot | Huawei (Petal Search) | Search | industry directories record it as robots-respecting; no dedicated official page found |
| YisouSpider | Alibaba (Shenma Search) | Search | official site exists (zhanzhang.sm.cn) but no token declaration found |
The most common blocking mistakes
The first one costs traffic every year:
This also blocks the search and user-triggered bots — you vanish from AI answers while training may continue anyway. Decide per category: training is your call; search and user-triggered are the ones you want.
User agents can be spoofed. Scanners routinely impersonate ChatGPT/Claude crawlers and probe `/.env`, `/wp-config`, `/actuator`. Judge by behaviour (do they fetch robots/sitemap/normal pages?).
Seeing `/xxxai/*` in a robots.txt block is a directory, not a crawler token.
Search and user-triggered bots usually drive the volume. Blocking training bots saves little and may cost you "being known" by the models.
How to verify identity (hardest to softest)
Google, OpenAI, Perplexity, Apple and Bing publish machine-readable IP range JSON; 360 publishes six ranges on its help page. A match confirms identity.
Reverse-resolve the IP; the hostname should belong to the vendor (Googlebot → *.googlebot.com, ByteDance → *.byted.org, Common Crawl → *.commoncrawl.org). ⚠️ 360 states it does not support nslookup — IP ranges only.
Genuine crawlers read `robots.txt`, `sitemap.xml` and normal pages; scanners ask for `.env`, `wp-config`, `actuator`. This layer is the fastest way to spot impostors.
FAQ
Not necessarily. ChatGPT’s live browsing uses ChatGPT-User and its search index uses OAI-SearchBot — they are independent from GPTBot. Blocking GPTBot alone usually does not affect live citations, but opts you out of training data.
Yes. Chinese AI search and assistants are major entry points for Chinese-language users. Most publish no IP ranges, so welcoming them on the content side (robots allow + structured content) is the pragmatic route.
Some vendors in China (e.g. DeepSeek, Zhipu) publish no crawler page or IP ranges. We observe and cross-check, never guess — such entries are explicitly labelled.
Write robots.txt policy, identify AI crawlers and scanners in your logs, and decide whether a self-declared crawler is genuine. Never block on user agent alone — blocking a real crawler hurts indexing and AI citations.


Bulk orders · integration help
Support only — not investment advice.
If you want to hand data to AI Agents rather than only being crawled: