Methodology

How the crawler access report works

AI Crawler Checker fetches the submitted public page, then uses its final URL and origin to inspect robots.txt and discovery files. Each bot row evaluates a documented user-agent token against the exact path and query. We use our own AI-Crawler-Checker/0.1 user agent for HTTP requests.

Three different kinds of evidence

  • Policy: an Allow or Disallow result from the robots.txt fetched by this scan.
  • HTTP access: what our server received. Official bots may see a different CDN, WAF, location, or authentication response.
  • Visits and use: bot identity needs trusted client-IP/log evidence. Indexing, training use, and citations need separate evidence; this scan does not verify them.

Rules and discovery

Robots matching and sitemap discovery

Rule matching handles UTF-8 and percent encoding, case-sensitive paths, wildcard and end-anchor rules, and Allow on equivalent ties. Encoding follows RFC 9309. Precedence uses the longest normalized pattern, including wildcard and end-anchor characters and retained percent escapes, following Google’s documented rule ordering. Specific user-agent groups take priority over the wildcard group; equally specific groups are combined. Individual crawlers may cache a different policy or interpret rules differently.

Sitemap discovery checks up to three distinct, public, absolute Sitemap declarations from usable robots.txt, plus the /sitemap.xml fallback. The report shows the chosen URL and checked candidates. We recognize sitemap or sitemap-index roots, but do not validate the full XML schema, decompress sitemap archives, or crawl child sitemaps and listed URLs.

JSON-LD and optional llms.txt

JSON-LD scripts are parsed as JSON and checked for a document object or nonempty array of objects. A pass means syntax and top-level shape passed, not that JSON-LD semantics, schema.org vocabulary, required properties, rich-result eligibility, or agreement with visible content were validated. Use a dedicated schema validator for those checks. The scanner reads initial HTML and does not execute page JavaScript.

llms.txt and llms-full.txt are optional experiments. They are excluded from general access, AEO, and visibility scores; missing JSON-LD is also unscored. The dedicated llms.txt tool grades the file’s basic Markdown structure, without claiming adoption or validating every linked page. Google’s AI search guidance requires no special AI text file or special structured data.

What statuses and scores mean

Use warnings and failures to prioritize review. A deliberate training opt-out can be a valid owner choice. Scores summarize the checklist; they do not predict ranking or citation.

Scoring weights and excluded checks

The general report prioritizes failed HTTP/indexing checks and blocked AI-search policies, then warnings and unknown observations. A deliberate training opt-out can be a valid owner choice. AEO and visibility checklist scores average 100 points for pass, 55 for warn, 35 for unknown, and 0 for fail, excluding rows marked “Not scored.” Bot-focused reports show the focused bot policy and page checks separately. A short text threshold is only a review heuristic, not a search-engine minimum word count.

These weights are this project’s prioritization choices. They are not Google eligibility rules, measured ranking factors, or a probability of appearing in AI answers. The AEO workflowsupports page editing; the visibility workflow explains how to collect actual citation observations after checking public access.

How non-standard fetchers are handled

Some user agents represent a person asking an AI product to fetch a page, or a product validating a submitted landing page, not an automatic web crawler. When official documentation says robots.txt may not apply or is generally ignored, the report labels that row as not scored instead of mixing it into the crawler-policy pass rate.

Access boundary

The scanner only checks public URLs and public site settings. It does not bypass logins, paywalls, firewall rules, bot defenses, or private systems. Requests are limited in time, redirects, response size, and request frequency so the report stays focused on public crawlability signals.

Each request has a nine-second time limit, up to five redirects, and an approximately 850 KB response cap. Oversized, unavailable, or non-HTML content can leave evidence incomplete. See the bot registry and official sources, privacy policy, and About this service.