> Source: https://ashareapi.com/en/docs/ai-crawlers/  ·  Markdown version for LLMs / AI agents

AI crawler list

# AI Crawler List & Verification (2026)
 Who to allow, who retrieves, who is user-triggered — and how to verify identity with official IP ranges and reverse DNS (user agents can be spoofed).

## Start with the three categories
 A vendor often runs **several** crawlers by purpose — blocking the training bot does **not** affect its search or user-triggered bots. Mixing them up is the most common cause of accidental blocking:
 |
| | Category | What it does | Example tokens | Recommendation

| | Training | Collects content for model training | GPTBot · ClaudeBot · Google-Extended · CCBot · Applebot-Extended · Bytespider · meta-externalagent | **Your content policy** (we allow: public content wants to be known)

| | Search / retrieval | Builds an index so the assistant can **cite your pages** | OAI-SearchBot · PerplexityBot · Claude-SearchBot · MistralAI-Index | **Allow** — this is the direct doorway to being cited

| | User-triggered | Fetches in **real time** for a user’s question | ChatGPT-User · Perplexity-User · Claude-User · MistralAI-User | **Allow** — a user is asking; if you are unreachable you lose the moment

 Training
 Collects content for model training
 GPTBot · ClaudeBot · Google-Extended · CCBot · Applebot-Extended · Bytespider · meta-externalagent
 **Your content policy** (we allow: public content wants to be known)

 Search / retrieval
 Builds an index so the assistant can **cite your pages**
 OAI-SearchBot · PerplexityBot · Claude-SearchBot · MistralAI-Index
 **Allow** — this is the direct doorway to being cited

 User-triggered
 Fetches in **real time** for a user’s question
 ChatGPT-User · Perplexity-User · Claude-User · MistralAI-User
 **Allow** — a user is asking; if you are unreachable you lose the moment

## The list (with official sources)
 Every token below is tied to an **official source**; where none exists we say so plainly (**observed only, not officially confirmed**) instead of guessing.
 |
| | token / user agent | Operator | Category | Official source

| | GPTBot | OpenAI | Training | openai.com/gptbot.json (IP ranges)

| | OAI-SearchBot | OpenAI | Search | openai.com/searchbot.json

| | ChatGPT-User | OpenAI | User-triggered | openai.com/chatgpt-user.json

| | ClaudeBot | Anthropic | Training | claude.com/crawling/bots.json (IP ranges)

| | Claude-SearchBot | Anthropic | Search | same (three-bot framework documented)

| | Claude-User | Anthropic | User-triggered | same

| | PerplexityBot | Perplexity | Search | perplexity.ai/perplexitybot.json

| | Perplexity-User | Perplexity | User-triggered | perplexity.ai/perplexity-user.json

| | Googlebot | Google | Search | developers.google.com/static/crawling/ipranges/common-crawlers.json

| | Google-Extended | Google | Training (token) | same — training only, **does not affect Search ranking**

| | Google-Agent | Google | User-triggered | same (added 2026-03)

| | Bingbot | Microsoft | Search (**also powers Copilot**) | bing.com/toolbox/bingbot.json

| | Applebot / Applebot-Extended | Apple | Search / training token | search.developer.apple.com/applebot.json

| | Amazonbot / Amzn-SearchBot / Amzn-User | Amazon | Training / Search / User-triggered | developer.amazon.com/amazonbot (rDNS is Amazon’s recommended check)

| | CCBot | Common Crawl | Training (open corpus) | official convention: rDNS to *.commoncrawl.org

| | meta-externalagent / Meta-ExternalFetcher | Meta | Training / user-triggered | developers.facebook.com/docs/sharing/webmasters/crawler

| | Bytespider | ByteDance (Doubao / TikTok search) | Training | zhanzhang.toutiao.com/docs/intro/26899 (official; rDNS suffix .byted.org)

| | DoubaoBot | ByteDance (Doubao) | AI Agent | listed in industry directories; **no dedicated official page found** (distinct from Bytespider)

 GPTBot
 OpenAI · Training
 openai.com/gptbot.json (IP ranges)

 OAI-SearchBot
 OpenAI · Search
 openai.com/searchbot.json

 ChatGPT-User
 OpenAI · User-triggered
 openai.com/chatgpt-user.json

 ClaudeBot
 Anthropic · Training
 claude.com/crawling/bots.json (IP ranges)

 Claude-SearchBot
 Anthropic · Search
 same (three-bot framework documented)

 Claude-User
 Anthropic · User-triggered
 same

 PerplexityBot
 Perplexity · Search
 perplexity.ai/perplexitybot.json

 Perplexity-User
 Perplexity · User-triggered
 perplexity.ai/perplexity-user.json

 Googlebot
 Google · Search
 developers.google.com/static/crawling/ipranges/common-crawlers.json

 Google-Extended
 Google · Training (token)
 same — training only, **does not affect Search ranking**

 Google-Agent
 Google · User-triggered
 same (added 2026-03)

 Bingbot
 Microsoft · Search (**also powers Copilot**)
 bing.com/toolbox/bingbot.json

 Applebot / Applebot-Extended
 Apple · Search / training token
 search.developer.apple.com/applebot.json

 Amazonbot / Amzn-SearchBot / Amzn-User
 Amazon · Training / Search / User-triggered
 developer.amazon.com/amazonbot (rDNS is Amazon’s recommended check)

 CCBot
 Common Crawl · Training (open corpus)
 official convention: rDNS to *.commoncrawl.org

 meta-externalagent / Meta-ExternalFetcher
 Meta · Training / user-triggered
 developers.facebook.com/docs/sharing/webmasters/crawler

 Bytespider
 ByteDance (Doubao / TikTok search) · Training
 zhanzhang.toutiao.com/docs/intro/26899 (official; rDNS suffix .byted.org)

 DoubaoBot
 ByteDance (Doubao) · AI Agent
 listed in industry directories; **no dedicated official page found** (distinct from Bytespider)

## China-based AI crawlers
 Chinese vendors publish less than their international peers (often no IP-range JSON). Official sources are marked; where there is none, we say so.
 |
| | token / user agent | Operator | Category | Official source

| | Baiduspider / Baiduspider-render | Baidu | Search / rendering | www.baidu.com/robots.txt (official declaration)

| | Sogou web spider / Sogou inst spider | Sogou (Tencent) | Search | www.sogou.com/robots.txt (official declaration)

| | 360Spider | 360 | Search | www.so.com/help/spider_ip.html — **six official IP ranges** published, with a note asking sites not to block it; officially states it **does not support nslookup** (IP ranges only)

| | Kimi-SearchBot | Moonshot AI (Kimi) | Search | kimi.ai/policies/kimi-crawlers (official policy page)

| | DeepSeekBot | DeepSeek | Training | no official crawler page found; industry directories record it as **not honoring robots.txt**

| | ChatGLM-Spider | Zhipu | Training / scraping | no official page found (industry entry is also incomplete)

| | QwenBot / TongyiBot / Qwen-User | Alibaba (Qwen) | Training / Search / User-triggered | sparse official English docs; recorded in industry directories

| | PetalBot | Huawei (Petal Search) | Search | industry directories record it as robots-respecting; no dedicated official page found

| | YisouSpider | Alibaba (Shenma Search) | Search | official site exists (zhanzhang.sm.cn) but no token declaration found

 Baiduspider / Baiduspider-render
 Baidu · Search / rendering
 www.baidu.com/robots.txt (official declaration)

 Sogou web spider / Sogou inst spider
 Sogou (Tencent) · Search
 www.sogou.com/robots.txt (official declaration)

 360Spider
 360 · Search
 www.so.com/help/spider_ip.html — **six official IP ranges** published, with a note asking sites not to block it; officially states it **does not support nslookup** (IP ranges only)

 Kimi-SearchBot
 Moonshot AI (Kimi) · Search
 kimi.ai/policies/kimi-crawlers (official policy page)

 DeepSeekBot
 DeepSeek · Training
 no official crawler page found; industry directories record it as **not honoring robots.txt**

 ChatGLM-Spider
 Zhipu · Training / scraping
 no official page found (industry entry is also incomplete)

 QwenBot / TongyiBot / Qwen-User
 Alibaba (Qwen) · Training / Search / User-triggered
 sparse official English docs; recorded in industry directories

 PetalBot
 Huawei (Petal Search) · Search
 industry directories record it as robots-respecting; no dedicated official page found

 YisouSpider
 Alibaba (Shenma Search) · Search
 official site exists (zhanzhang.sm.cn) but no token declaration found

## The most common blocking mistakes
 The first one costs traffic every year:
 Blocking everything with `Disallow: /` to "stop AI training"
 This also blocks the **search and user-triggered** bots — you vanish from AI answers while training may continue anyway. Decide **per category**: training is your call; search and user-triggered are the ones you want.

 Trusting the user agent
 **User agents can be spoofed.** Scanners routinely impersonate ChatGPT/Claude crawlers and probe `/.env`, `/wp-config`, `/actuator`. Judge by **behaviour** (do they fetch robots/sitemap/normal pages?).

 Mistaking a path for a declared bot
 Seeing `/xxxai/*` in a robots.txt block is a **directory**, not a crawler token.

 Thinking blocking training bots saves traffic
 Search and user-triggered bots usually drive the volume. Blocking training bots saves little and may cost you "being known" by the models.

## How to verify identity (hardest to softest)
 ① Official IP ranges (hardest)
 Google, OpenAI, Perplexity, Apple and Bing publish **machine-readable IP range JSON**; 360 publishes six ranges on its help page. A match confirms identity.

 ② Reverse DNS
 Reverse-resolve the IP; the hostname should belong to the vendor (Googlebot → *.googlebot.com, ByteDance → *.byted.org, Common Crawl → *.commoncrawl.org). ⚠️ 360 states it does not support nslookup — IP ranges only.

 ③ Behaviour
 Genuine crawlers read `robots.txt`, `sitemap.xml` and normal pages; scanners ask for `.env`, `wp-config`, `actuator`. **This layer is the fastest way to spot impostors.**

## FAQ
 If I block GPTBot, does ChatGPT stop seeing my site?
 Not necessarily. ChatGPT’s **live browsing** uses ChatGPT-User and its **search index** uses OAI-SearchBot — they are independent from GPTBot. Blocking GPTBot alone usually does not affect live citations, but opts you out of training data.

 Should I allow China-based AI crawlers?
 Yes. Chinese AI search and assistants are major entry points for Chinese-language users. Most publish no IP ranges, so welcoming them on the content side (robots allow + structured content) is the pragmatic route.

 Why are some crawlers marked "not officially confirmed"?
 Some vendors in China (e.g. DeepSeek, Zhipu) publish no crawler page or IP ranges. We **observe and cross-check**, never guess — such entries are explicitly labelled.

 What can I do with this list?
 Write robots.txt policy, identify AI crawlers and scanners in your logs, and decide whether a self-declared crawler is genuine. **Never block on user agent alone** — blocking a real crawler hurts indexing and AI citations.

 ![Buy on Xianyu]() Buy on Xianyu

 ![Scan on WeChat]() Scan on WeChat

 Purchase & support
 Bulk orders · integration help
 Support only — **not investment advice**.

 Related
 If you want to hand data to AI Agents rather than only being crawled:

 [MCP setup guide](/en/mcp)[China platform guide](/en/docs/mcp-china)
