ClaudeBot is not every Claude request
ClaudeBot, Claude-SearchBot, and Claude-User are shown separately so training, search retrieval, and user-triggered access do not collapse into one status.
Anthropic crawler policy
Run a focused ClaudeBot access report and compare Anthropic crawler roles in one place. The checker separates ClaudeBot from Claude-SearchBot and Claude-User, then checks whether page-level signals make the URL usable.
Rules, detection, and verification
Set your crawl policy, troubleshoot blocked requests, and verify bot traffic in your logs.
Anthropic distinguishes training (ClaudeBot), search (Claude-SearchBot), and user requests (Claude-User). Its documentation says all three honor robots.txt. This example blocks training while allowing search and user-directed retrieval. Read the operator’s crawler documentation.
User-agent: ClaudeBot
Disallow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /Merge the groups you want into the existing root /robots.txt; preserve unrelated policies and sitemap declarations. Test a representative public page and a path that should stay restricted. Each hostname needs its own policy. To permit training instead, use Allow: / in the ClaudeBot group and review any more specific Disallow rules. robots.txt is not authentication for private content.
Dashboard controls and defaults can change. Use Cloudflare’s current AI bot policy instructions. An HTTP 200 from this tool only describes its own request; official bots can receive a different response.
Filter your access logs for ClaudeBot in the user-agent field. Retain the client IP, time, method, path, status, and bytes served. A matching string is a claim anyone can send. Compare the actual client IP with the operator’s current published IP ranges using CIDR membership, not text-prefix matching. Behind a proxy, use its authenticated client-IP field; do not trust an arbitrary forwarded header from a public request.
Save this script, then run python verify_bot_ip.py YOUR_LOG_IP. It downloads the public range list; the log IP stays in your local process. An unmatched result needs investigation and is not conclusive proof of impersonation.
# Save as verify_bot_ip.py; use a client IP from your trusted edge logs.
import ipaddress, json, sys, urllib.request
address = ipaddress.ip_address(sys.argv[1])
with urllib.request.urlopen("https://claude.com/crawling/bots.json", timeout=10) as response:
prefixes = json.load(response)["prefixes"]
ranges = [ipaddress.ip_network(value) for item in prefixes
for key, value in item.items() if key in ("ipv4Prefix", "ipv6Prefix")]
print("In published ranges" if any(address in network for network in ranges)
else "Not in current published ranges")Anthropic’s published list identifies its crawler network; it does not distinguish which Claude role made an individual request. Keep the user-agent and request details alongside the IP evidence. A verified request with a successful response supports a visit claim. It does not establish training use, indexing, or citation.
Run this URL check after publishing the rule, copy the report with its timestamp, and compare it with later edge logs. If the rule now permits access but requests still fail, investigate HTTP or WAF behavior. If requests succeed but no answer cites your page, continue with the visibility measurement checklist.
Included checks
Review crawler rules, page responses, and discovery files for the URL you submit.
ClaudeBot, Claude-SearchBot, and Claude-User are shown separately so training, search retrieval, and user-triggered access do not collapse into one status.
The report evaluates the tested URL path, not only the homepage, and shows the matching robots.txt rule line when available.
Even when ClaudeBot is allowed, noindex headers, missing sitemap signals, or thin readable HTML can reduce the practical value of access.
Methodology
Each recommendation ties back to a public robots.txt rule, header, meta tag, file, or fetched page signal.
The focused report checks ClaudeBot first, then shows Claude-SearchBot and Claude-User for role clarity.
The scanner parses user-agent groups and Allow/Disallow rules for the submitted URL path.
The result includes page fetch status, noindex directives, sitemap, and llms.txt because bot access alone is not a visibility guarantee.
Related tools
Move from one bot decision into broader robots, visibility, and crawler-readiness checks.
FAQ
Short answers for site owners deciding how to handle crawler-specific robots.txt policies.
It checks whether ClaudeBot is allowed or blocked by robots.txt for the exact submitted path, then reviews related Claude agents and page-level blockers such as noindex headers, sitemap, and llms.txt.
No. This page separates ClaudeBot from Claude-SearchBot and Claude-User so training, search/retrieval, and user-triggered behavior can be reviewed independently.
robots.txt access is only one layer. A page can be allowed for a crawler but still carry noindex directives or weak discovery signals.
No. The scanner does not create public saved reports.