OpenAI crawler policy

GPTBot Checker

Run a focused GPTBot access report for the exact public URL path you care about. The checker separates GPTBot from OAI-SearchBot, OAI-AdsBot, and ChatGPT-User so training, search retrieval, ad validation, and user-triggered access do not get blurred together.

Rules, detection, and verification

How to control and verify GPTBot

Set your crawl policy, troubleshoot blocked requests, and verify bot traffic in your logs.

1. Choose training and search policies separately

OpenAI assigns training to GPTBot and search to OAI-SearchBot. The example opts out of the former while permitting the latter. ChatGPT-User handles user requests, where robots.txt may not apply. Read the operator’s crawler documentation.

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

Merge the groups you want into the existing root /robots.txt; preserve unrelated policies and sitemap declarations. Test a representative public page and a path that should stay restricted. Each hostname needs its own policy. To permit training instead, use Allow: / in the GPTBot group and review any more specific Disallow rules. robots.txt is not authentication for private content.

2. Investigate a Cloudflare block

  1. Record the failing URL, timestamp, HTTP status, and request or Ray ID. Compare the exact host and path with the checker’s matched rule.
  2. In Cloudflare Security Settings, review the AI bot policies for Training, Search, and Agent activity, plus any matching custom WAF rules. Keep the categories aligned with your intended robots policy.
  3. Inspect the corresponding security event. Change only the rule responsible for an unwanted block, then check a new request. A broad WAF bypass based on a user-agent string also admits spoofed traffic.

Dashboard controls and defaults can change. Use Cloudflare’s current AI bot policy instructions. An HTTP 200 from this tool only describes its own request; official bots can receive a different response.

3. Detect a claimed bot, then verify its source IP

Filter your access logs for GPTBot in the user-agent field. Retain the client IP, time, method, path, status, and bytes served. A matching string is a claim anyone can send. Compare the actual client IP with the operator’s current published IP ranges using CIDR membership, not text-prefix matching. Behind a proxy, use its authenticated client-IP field; do not trust an arbitrary forwarded header from a public request.

Check one log IP locally with Python

Save this script, then run python verify_bot_ip.py YOUR_LOG_IP. It downloads the public range list; the log IP stays in your local process. An unmatched result needs investigation and is not conclusive proof of impersonation.

# Save as verify_bot_ip.py; use a client IP from your trusted edge logs.
import ipaddress, json, sys, urllib.request

address = ipaddress.ip_address(sys.argv[1])
with urllib.request.urlopen("https://openai.com/gptbot.json", timeout=10) as response:
    prefixes = json.load(response)["prefixes"]
ranges = [ipaddress.ip_network(value) for item in prefixes
          for key, value in item.items() if key in ("ipv4Prefix", "ipv6Prefix")]
print("In published ranges" if any(address in network for network in ranges)
      else "Not in current published ranges")

Use the separate OAI-SearchBot range list when investigating OpenAI search traffic; GPTBot’s list verifies the training crawler’s network. A verified request with a successful response supports a visit claim. It does not establish training use, indexing, or citation.

4. Retest and keep the evidence

Run this URL check after publishing the rule, copy the report with its timestamp, and compare it with later edge logs. If the rule now permits access but requests still fail, investigate HTTP or WAF behavior. If requests succeed but no answer cites your page, continue with the visibility measurement checklist.

Included checks

What you can check

Review crawler rules, page responses, and discovery files for the URL you submit.

GPTBot is the training crawler

GPTBot is treated as an OpenAI training/model-improvement crawler in this report, not as the same role as search retrieval or user-triggered fetches.

Path-specific evidence

The result evaluates robots.txt against the submitted URL path and surfaces the matching rule line when a rule applies.

OpenAI roles compared

The report shows OAI-SearchBot, OAI-AdsBot, and ChatGPT-User next to GPTBot so you can separate search retrieval, ad validation, user-triggered access, and training-policy decisions.

Methodology

How the bot decision is made

Each recommendation ties back to a public robots.txt rule, header, meta tag, file, or fetched page signal.

View rule matching and report details

Fetch public robots.txt

The scanner fetches same-origin /robots.txt with bounded redirects and evaluates the exact submitted path.

Prioritize GPTBot

The focused report lifts GPTBot above the full bot table and includes related OpenAI agents for policy contrast.

Check page blockers

The scan also reviews HTTP status, meta robots, X-Robots-Tag, sitemap, and llms.txt because access is only useful when the page can be discovered and read.

Scope of the report

Scope of the report
  • This checker does not prove whether OpenAI has used, will use, or will cite the submitted page.
  • robots.txt is a public policy signal; individual crawler behavior can change and should be verified against official documentation.
  • The scan does not log in, bypass bot defenses, fetch private URLs, or publish public scan reports.

Related tools

Move from one bot decision into broader robots, visibility, and crawler-readiness checks.

FAQ

Common questions

Short answers for site owners deciding how to handle crawler-specific robots.txt policies.

What does the GPTBot checker test?

It checks whether GPTBot is allowed or blocked by robots.txt for the exact submitted path, then adds page-level blockers such as noindex headers, sitemap availability, and llms.txt context.

Is GPTBot the same as OAI-SearchBot or ChatGPT-User?

No. The page treats GPTBot as a training crawler, OAI-SearchBot as search/retrieval, OAI-AdsBot as ad validation, and ChatGPT-User as user-triggered context so site owners can make separate policy choices.

Can I block GPTBot but allow OpenAI search retrieval?

Yes, robots.txt can express different rules for different user agents. The focused report helps confirm whether the intended split applies to the tested path.

Do you store GPTBot scan results?

No. Scans are processed for the response and are not saved as public reports.