Anthropic crawler policy

ClaudeBot Checker

Run a focused ClaudeBot access report and compare Anthropic crawler roles in one place. The checker separates ClaudeBot from Claude-SearchBot and Claude-User, then checks whether page-level signals make the URL usable.

Rules, detection, and verification

How to control and verify ClaudeBot

Set your crawl policy, troubleshoot blocked requests, and verify bot traffic in your logs.

1. Choose training and search policies separately

Anthropic distinguishes training (ClaudeBot), search (Claude-SearchBot), and user requests (Claude-User). Its documentation says all three honor robots.txt. This example blocks training while allowing search and user-directed retrieval. Read the operator’s crawler documentation.

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

Merge the groups you want into the existing root /robots.txt; preserve unrelated policies and sitemap declarations. Test a representative public page and a path that should stay restricted. Each hostname needs its own policy. To permit training instead, use Allow: / in the ClaudeBot group and review any more specific Disallow rules. robots.txt is not authentication for private content.

2. Investigate a Cloudflare block

  1. Record the failing URL, timestamp, HTTP status, and request or Ray ID. Compare the exact host and path with the checker’s matched rule.
  2. In Cloudflare Security Settings, review the AI bot policies for Training, Search, and Agent activity, plus any matching custom WAF rules. Keep the categories aligned with your intended robots policy.
  3. Inspect the corresponding security event. Change only the rule responsible for an unwanted block, then check a new request. A broad WAF bypass based on a user-agent string also admits spoofed traffic.

Dashboard controls and defaults can change. Use Cloudflare’s current AI bot policy instructions. An HTTP 200 from this tool only describes its own request; official bots can receive a different response.

3. Detect a claimed bot, then verify its source IP

Filter your access logs for ClaudeBot in the user-agent field. Retain the client IP, time, method, path, status, and bytes served. A matching string is a claim anyone can send. Compare the actual client IP with the operator’s current published IP ranges using CIDR membership, not text-prefix matching. Behind a proxy, use its authenticated client-IP field; do not trust an arbitrary forwarded header from a public request.

Check one log IP locally with Python

Save this script, then run python verify_bot_ip.py YOUR_LOG_IP. It downloads the public range list; the log IP stays in your local process. An unmatched result needs investigation and is not conclusive proof of impersonation.

# Save as verify_bot_ip.py; use a client IP from your trusted edge logs.
import ipaddress, json, sys, urllib.request

address = ipaddress.ip_address(sys.argv[1])
with urllib.request.urlopen("https://claude.com/crawling/bots.json", timeout=10) as response:
    prefixes = json.load(response)["prefixes"]
ranges = [ipaddress.ip_network(value) for item in prefixes
          for key, value in item.items() if key in ("ipv4Prefix", "ipv6Prefix")]
print("In published ranges" if any(address in network for network in ranges)
      else "Not in current published ranges")

Anthropic’s published list identifies its crawler network; it does not distinguish which Claude role made an individual request. Keep the user-agent and request details alongside the IP evidence. A verified request with a successful response supports a visit claim. It does not establish training use, indexing, or citation.

4. Retest and keep the evidence

Run this URL check after publishing the rule, copy the report with its timestamp, and compare it with later edge logs. If the rule now permits access but requests still fail, investigate HTTP or WAF behavior. If requests succeed but no answer cites your page, continue with the visibility measurement checklist.

Included checks

What you can check

Review crawler rules, page responses, and discovery files for the URL you submit.

ClaudeBot is not every Claude request

ClaudeBot, Claude-SearchBot, and Claude-User are shown separately so training, search retrieval, and user-triggered access do not collapse into one status.

Robots rules are path-specific

The report evaluates the tested URL path, not only the homepage, and shows the matching robots.txt rule line when available.

Discovery blockers still matter

Even when ClaudeBot is allowed, noindex headers, missing sitemap signals, or thin readable HTML can reduce the practical value of access.

Methodology

How the bot decision is made

Each recommendation ties back to a public robots.txt rule, header, meta tag, file, or fetched page signal.

View rule matching and report details

Evaluate Anthropic agents separately

The focused report checks ClaudeBot first, then shows Claude-SearchBot and Claude-User for role clarity.

Use exact robots matching

The scanner parses user-agent groups and Allow/Disallow rules for the submitted URL path.

Add page-readiness context

The result includes page fetch status, noindex directives, sitemap, and llms.txt because bot access alone is not a visibility guarantee.

Scope of the report

Scope of the report
  • This checker does not prove whether Anthropic has used, will use, or will cite the submitted page.
  • Anthropic documents robots.txt support for Claude-User as well as ClaudeBot and Claude-SearchBot; their purposes still differ.
  • The scan does not bypass login, WAF rules, paywalls, or private address protections.

Related tools

Move from one bot decision into broader robots, visibility, and crawler-readiness checks.

FAQ

Common questions

Short answers for site owners deciding how to handle crawler-specific robots.txt policies.

What does the ClaudeBot checker test?

It checks whether ClaudeBot is allowed or blocked by robots.txt for the exact submitted path, then reviews related Claude agents and page-level blockers such as noindex headers, sitemap, and llms.txt.

Is ClaudeBot the same as Claude-SearchBot or Claude-User?

No. This page separates ClaudeBot from Claude-SearchBot and Claude-User so training, search/retrieval, and user-triggered behavior can be reviewed independently.

Why does the page check headers and metadata too?

robots.txt access is only one layer. A page can be allowed for a crawler but still carry noindex directives or weak discovery signals.

Do you store ClaudeBot scan results?

No. The scanner does not create public saved reports.