AI crawlers explained: GPTBot, ClaudeBot, PerplexityBot and how to see them
Which AI crawlers visit your website, what each one is for, why your analytics does not show them, and how to measure them from your server's access log.
Every AI assistant that answers questions about the web depends on crawlers: programs that fetch pages so they can be used for training, for search results, or to answer a specific user's question. If your website is public, these crawlers are almost certainly visiting it already.
You will not find them in a typical analytics dashboard. This article explains which crawlers exist, what each one does, and how to see them.
Three kinds of AI crawler
It helps to separate AI crawlers by purpose, because the same company often runs several:
- Training crawlers collect pages to train language models.
- Search crawlers build an index that an assistant searches when it answers with sources.
- User-triggered fetchers load a specific page because a user asked the assistant to look at it, or because the assistant decided to open a link while answering.
The distinction matters if you ever decide to block one: blocking a training crawler and blocking a search crawler have very different effects on whether assistants can recommend you.
The crawlers you are most likely to see
| User agent | Operator | Purpose |
|---|---|---|
| GPTBot | OpenAI | Training |
| OAI-SearchBot | OpenAI | Search results in ChatGPT |
| ChatGPT-User | OpenAI | Pages fetched on behalf of a ChatGPT user |
| ClaudeBot | Anthropic | Training |
| Claude-SearchBot | Anthropic | Search |
| Claude-User | Anthropic | Pages fetched on behalf of a Claude user |
| PerplexityBot | Perplexity | Search index |
| Perplexity-User | Perplexity | Pages fetched on behalf of a user |
| CCBot | Common Crawl | Open web archive widely used for AI training |
| Bytespider | ByteDance | Crawling for ByteDance services |
| Amazonbot | Amazon | Crawling for Amazon services, including assistants |
| Meta-ExternalAgent | Meta | Crawling for AI products |
Two names you may see in robots.txt guides are not separate crawlers at all. Google-Extended and Applebot-Extended are control tokens: Google and Apple crawl with their normal bots, and these tokens only tell them whether your content may be used for their AI models.
User agent strings can be faked, so treat the names above as a strong hint rather than proof. The major operators publish the IP ranges their crawlers use if you ever need to verify a specific request.
Why your analytics does not show them
Script-based analytics, including Google Analytics, Matomo's JavaScript tracker and ours, counts a visit when a small script runs in the visitor's browser. Crawlers download your HTML but generally do not execute JavaScript, so the script never runs and the visit is never counted.
That is mostly a good thing, because bots would otherwise inflate your visitor numbers. But it means the only reliable record of AI crawlers is on your server.
How to find them in your access log
Every request to your site is written to the web server's access log. On a typical Linux server with nginx it lives at /var/log/nginx/access.log, and on Apache at /var/log/apache2/access.log or a path set by your hosting panel. Each line includes the requested path, the status code and the user agent.
To list the AI crawler requests from a log:
grep -iE 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|Claude-SearchBot|PerplexityBot|Perplexity-User|CCBot|Bytespider|Amazonbot|Meta-ExternalAgent' /var/log/nginx/access.log
To count requests per crawler, pipe that through a short summary:
grep -oiE 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|PerplexityBot|CCBot|Bytespider' /var/log/nginx/access.log | sort | uniq -c | sort -rn
Look for three things:
- Which pages they read. Are your most important pages being fetched at all?
- Status codes. A crawler that keeps receiving 403 or 404 responses cannot use your content. Firewalls and bot protection sometimes block AI crawlers without anyone deciding to.
- Trends. A rise in user-triggered fetchers such as ChatGPT-User usually means people are asking assistants about topics you cover.
Should you block AI crawlers?
That is a business decision, not a technical one. Blocking training crawlers in robots.txt keeps your content out of future model training. Blocking search and user-triggered crawlers also keeps you out of answers that cite sources, which means assistants cannot recommend you or send you visitors.
A sensible order is to measure first: see which crawlers visit, which pages they read, and how much referral traffic the same assistants send you. Then decide with numbers in front of you.
If you do want to block a crawler, robots.txt uses the user agent token:
User-agent: GPTBot
Disallow: /
Also check your CDN or firewall settings. Some providers offer a switch that blocks AI crawlers for every site on the account, which can be turned on without anyone noticing.
Measuring crawlers alongside your visitors
Our analytics lets you upload an access log, or send it automatically every hour with one cron line, and counts only the AI crawler requests in it. IP addresses in the log are never stored, and sending the same lines twice never double-counts them. The results sit next to your AI referral traffic, so you can see both which pages AI crawlers read and which ones assistants actually send people to.
Learn more about AI traffic analytics or start a free trial.