GPTBot is the training crawler
GPTBot is treated as an OpenAI training/model-improvement crawler in this report, not as the same role as search retrieval or user-triggered fetches.
OpenAI crawler policy
Run a focused GPTBot access report for the exact public URL path you care about. The checker separates GPTBot from OAI-SearchBot, OAI-AdsBot, and ChatGPT-User so training, search retrieval, ad validation, and user-triggered access do not get blurred together.
Rules, detection, and verification
Set your crawl policy, troubleshoot blocked requests, and verify bot traffic in your logs.
OpenAI assigns training to GPTBot and search to OAI-SearchBot. The example opts out of the former while permitting the latter. ChatGPT-User handles user requests, where robots.txt may not apply. Read the operator’s crawler documentation.
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /Merge the groups you want into the existing root /robots.txt; preserve unrelated policies and sitemap declarations. Test a representative public page and a path that should stay restricted. Each hostname needs its own policy. To permit training instead, use Allow: / in the GPTBot group and review any more specific Disallow rules. robots.txt is not authentication for private content.
Dashboard controls and defaults can change. Use Cloudflare’s current AI bot policy instructions. An HTTP 200 from this tool only describes its own request; official bots can receive a different response.
Filter your access logs for GPTBot in the user-agent field. Retain the client IP, time, method, path, status, and bytes served. A matching string is a claim anyone can send. Compare the actual client IP with the operator’s current published IP ranges using CIDR membership, not text-prefix matching. Behind a proxy, use its authenticated client-IP field; do not trust an arbitrary forwarded header from a public request.
Save this script, then run python verify_bot_ip.py YOUR_LOG_IP. It downloads the public range list; the log IP stays in your local process. An unmatched result needs investigation and is not conclusive proof of impersonation.
# Save as verify_bot_ip.py; use a client IP from your trusted edge logs.
import ipaddress, json, sys, urllib.request
address = ipaddress.ip_address(sys.argv[1])
with urllib.request.urlopen("https://openai.com/gptbot.json", timeout=10) as response:
prefixes = json.load(response)["prefixes"]
ranges = [ipaddress.ip_network(value) for item in prefixes
for key, value in item.items() if key in ("ipv4Prefix", "ipv6Prefix")]
print("In published ranges" if any(address in network for network in ranges)
else "Not in current published ranges")Use the separate OAI-SearchBot range list when investigating OpenAI search traffic; GPTBot’s list verifies the training crawler’s network. A verified request with a successful response supports a visit claim. It does not establish training use, indexing, or citation.
Run this URL check after publishing the rule, copy the report with its timestamp, and compare it with later edge logs. If the rule now permits access but requests still fail, investigate HTTP or WAF behavior. If requests succeed but no answer cites your page, continue with the visibility measurement checklist.
Included checks
Review crawler rules, page responses, and discovery files for the URL you submit.
GPTBot is treated as an OpenAI training/model-improvement crawler in this report, not as the same role as search retrieval or user-triggered fetches.
The result evaluates robots.txt against the submitted URL path and surfaces the matching rule line when a rule applies.
The report shows OAI-SearchBot, OAI-AdsBot, and ChatGPT-User next to GPTBot so you can separate search retrieval, ad validation, user-triggered access, and training-policy decisions.
Methodology
Each recommendation ties back to a public robots.txt rule, header, meta tag, file, or fetched page signal.
The scanner fetches same-origin /robots.txt with bounded redirects and evaluates the exact submitted path.
The focused report lifts GPTBot above the full bot table and includes related OpenAI agents for policy contrast.
The scan also reviews HTTP status, meta robots, X-Robots-Tag, sitemap, and llms.txt because access is only useful when the page can be discovered and read.
Related tools
Move from one bot decision into broader robots, visibility, and crawler-readiness checks.
FAQ
Short answers for site owners deciding how to handle crawler-specific robots.txt policies.
It checks whether GPTBot is allowed or blocked by robots.txt for the exact submitted path, then adds page-level blockers such as noindex headers, sitemap availability, and llms.txt context.
No. The page treats GPTBot as a training crawler, OAI-SearchBot as search/retrieval, OAI-AdsBot as ad validation, and ChatGPT-User as user-triggered context so site owners can make separate policy choices.
Yes, robots.txt can express different rules for different user agents. The focused report helps confirm whether the intended split applies to the tested path.
No. Scans are processed for the response and are not saved as public reports.