AI crawler list

AI Crawler List & Verification (2026)

Who to allow, who retrieves, who is user-triggered — and how to verify identity with official IP ranges and reverse DNS (user agents can be spoofed).

Start with the three categories

A vendor often runs several crawlers by purpose — blocking the training bot does not affect its search or user-triggered bots. Mixing them up is the most common cause of accidental blocking:

Training

Collects content for model training

GPTBot · ClaudeBot · Google-Extended · CCBot · Applebot-Extended · Bytespider · meta-externalagent

Your content policy (we allow: public content wants to be known)

Search / retrieval

Builds an index so the assistant can cite your pages

OAI-SearchBot · PerplexityBot · Claude-SearchBot · MistralAI-Index

Allow — this is the direct doorway to being cited

User-triggered

Fetches in real time for a user’s question

ChatGPT-User · Perplexity-User · Claude-User · MistralAI-User

Allow — a user is asking; if you are unreachable you lose the moment

The list (with official sources)

Every token below is tied to an official source; where none exists we say so plainly (observed only, not officially confirmed) instead of guessing.

GPTBot
OpenAI · Training
openai.com/gptbot.json (IP ranges)
OAI-SearchBot
OpenAI · Search
openai.com/searchbot.json
ChatGPT-User
OpenAI · User-triggered
openai.com/chatgpt-user.json
ClaudeBot
Anthropic · Training
claude.com/crawling/bots.json (IP ranges)
Claude-SearchBot
Anthropic · Search
same (three-bot framework documented)
Claude-User
Anthropic · User-triggered
same
PerplexityBot
Perplexity · Search
perplexity.ai/perplexitybot.json
Perplexity-User
Perplexity · User-triggered
perplexity.ai/perplexity-user.json
Googlebot
Google · Search
developers.google.com/static/crawling/ipranges/common-crawlers.json
Google-Extended
Google · Training (token)
same — training only, does not affect Search ranking
Google-Agent
Google · User-triggered
same (added 2026-03)
Bingbot
Microsoft · Search (also powers Copilot)
bing.com/toolbox/bingbot.json
Applebot / Applebot-Extended
Apple · Search / training token
search.developer.apple.com/applebot.json
Amazonbot / Amzn-SearchBot / Amzn-User
Amazon · Training / Search / User-triggered
developer.amazon.com/amazonbot (rDNS is Amazon’s recommended check)
CCBot
Common Crawl · Training (open corpus)
official convention: rDNS to *.commoncrawl.org
meta-externalagent / Meta-ExternalFetcher
Meta · Training / user-triggered
developers.facebook.com/docs/sharing/webmasters/crawler
Bytespider
ByteDance (Doubao / TikTok search) · Training
zhanzhang.toutiao.com/docs/intro/26899 (official; rDNS suffix .byted.org)
DoubaoBot
ByteDance (Doubao) · AI Agent
listed in industry directories; no dedicated official page found (distinct from Bytespider)

China-based AI crawlers

Chinese vendors publish less than their international peers (often no IP-range JSON). Official sources are marked; where there is none, we say so.

Baiduspider / Baiduspider-render
Baidu · Search / rendering
www.baidu.com/robots.txt (official declaration)
Sogou web spider / Sogou inst spider
Sogou (Tencent) · Search
www.sogou.com/robots.txt (official declaration)
360Spider
360 · Search
www.so.com/help/spider_ip.html — six official IP ranges published, with a note asking sites not to block it; officially states it does not support nslookup (IP ranges only)
Kimi-SearchBot
Moonshot AI (Kimi) · Search
kimi.ai/policies/kimi-crawlers (official policy page)
DeepSeekBot
DeepSeek · Training
no official crawler page found; industry directories record it as not honoring robots.txt
ChatGLM-Spider
Zhipu · Training / scraping
no official page found (industry entry is also incomplete)
QwenBot / TongyiBot / Qwen-User
Alibaba (Qwen) · Training / Search / User-triggered
sparse official English docs; recorded in industry directories
PetalBot
Huawei (Petal Search) · Search
industry directories record it as robots-respecting; no dedicated official page found
YisouSpider
Alibaba (Shenma Search) · Search
official site exists (zhanzhang.sm.cn) but no token declaration found

The most common blocking mistakes

The first one costs traffic every year:

Blocking everything with `Disallow: /` to "stop AI training"

This also blocks the search and user-triggered bots — you vanish from AI answers while training may continue anyway. Decide per category: training is your call; search and user-triggered are the ones you want.

Trusting the user agent

User agents can be spoofed. Scanners routinely impersonate ChatGPT/Claude crawlers and probe `/.env`, `/wp-config`, `/actuator`. Judge by behaviour (do they fetch robots/sitemap/normal pages?).

Mistaking a path for a declared bot

Seeing `/xxxai/*` in a robots.txt block is a directory, not a crawler token.

Thinking blocking training bots saves traffic

Search and user-triggered bots usually drive the volume. Blocking training bots saves little and may cost you "being known" by the models.

How to verify identity (hardest to softest)

① Official IP ranges (hardest)

Google, OpenAI, Perplexity, Apple and Bing publish machine-readable IP range JSON; 360 publishes six ranges on its help page. A match confirms identity.

② Reverse DNS

Reverse-resolve the IP; the hostname should belong to the vendor (Googlebot → *.googlebot.com, ByteDance → *.byted.org, Common Crawl → *.commoncrawl.org). ⚠️ 360 states it does not support nslookup — IP ranges only.

③ Behaviour

Genuine crawlers read `robots.txt`, `sitemap.xml` and normal pages; scanners ask for `.env`, `wp-config`, `actuator`. This layer is the fastest way to spot impostors.

FAQ

If I block GPTBot, does ChatGPT stop seeing my site?

Not necessarily. ChatGPT’s live browsing uses ChatGPT-User and its search index uses OAI-SearchBot — they are independent from GPTBot. Blocking GPTBot alone usually does not affect live citations, but opts you out of training data.

Should I allow China-based AI crawlers?

Yes. Chinese AI search and assistants are major entry points for Chinese-language users. Most publish no IP ranges, so welcoming them on the content side (robots allow + structured content) is the pragmatic route.

Why are some crawlers marked "not officially confirmed"?

Some vendors in China (e.g. DeepSeek, Zhipu) publish no crawler page or IP ranges. We observe and cross-check, never guess — such entries are explicitly labelled.

What can I do with this list?

Write robots.txt policy, identify AI crawlers and scanners in your logs, and decide whether a self-declared crawler is genuine. Never block on user agent alone — blocking a real crawler hurts indexing and AI citations.

Buy on Xianyu
Buy on Xianyu
Scan on WeChat
Scan on WeChat
Purchase & support

Bulk orders · integration help

Support only — not investment advice.

Related

If you want to hand data to AI Agents rather than only being crawled: