AI Citation Readiness Checker

Audit your robots.txt, llms.txt, and page HTML for AI-crawler access and citation-readiness. Three evidence-based layers, instant results, nothing uploaded.

Client-side only Nothing stored No account needed
Updated for 2026 Last verified July 2026

This is the file AI crawlers read most. The audit runs entirely in this browser; nothing is uploaded.

One representative page is enough. The checker reads structure, not your whole site.

This tool runs locally in your browser.

Output will appear here.
Ask an AI to explain this result
Share this result

Results are estimates based on your input only; entered values and result contents are not stored by BoringToolsKit.

Your inputs and results stay in this browser. This page may send privacy-bounded aggregate interaction events described in Privacy; it does not send your entered values or result contents.

Use this result

Share the current inputs or ask ChatGPT to explain the calculation in context.

More options
Report a calculation issue
Direct answer

What does this calculator estimate?

The AI Citation Readiness Checker audits your robots.txt, llms.txt, and one page of HTML for the structure AI crawlers and answer engines actually read. It scores three evidence-based layers: crawler access, page extractability, and llms.txt quality. Everything runs in your browser; nothing is uploaded.

  • robots.txt and sitemap.xml are the only infrastructure paths every major AI crawler reads daily
  • AI-facing JSON endpoints (/ai/summary.json, ai.txt) received zero crawler requests in published 15-day logs
  • On BoringToolsKit's own site, robots.txt was fetched 845 times in 7 days vs llms.txt 24 times
  • FAQPage schema affects how AI quotes a page, not whether crawlers visit it
  • Citation-driven AI visits are growing 3.3x faster than crawler visits (173-site panel, mid-2026)

The three layers AI visibility actually runs on

Access: robots.txt is the most-fetched file on any site, read almost daily by every major AI crawler. A single Disallow line for GPTBot or ClaudeBot makes you invisible to that assistant by configuration, and most site owners never notice. Orientation: sitemap.xml tells crawlers which pages matter; ClaudeBot reads it daily. Entity: the citation layer — a one-sentence answer near the top of the page, JSON-LD structured data, FAQPage markup, and honest dateModified dates give AI engines clean facts to extract and quote.

What the crawl-log evidence says

Two independent log studies agree with each other and with our own production data. A 15-day Cloudflare Worker study (geo010.com) found robots.txt and sitemap.xml were the only paths every major AI crawler used, while purpose-built AI discovery endpoints received zero requests. An 83-site study watched OpenAI's crawler read robots.txt 3,990 times and llms.txt 7 times in 12 weeks. On our own 854-tool site, the 7-day ratio was 845 robots.txt fetches to 24 llms.txt fetches. The boring files win.

Why there is no AI JSON endpoint to build

A common vendor pitch is adding /ai/summary.json or /.well-known/ai.txt so 'AI engines can read your site'. The evidence contradicts this: geo010 logged zero requests to any AI-facing endpoint across 15 days, across GPTBot, ClaudeBot, PerplexityBot, and OAI-SearchBot. Crawlers find content through robots.txt, sitemaps, and links. This checker encodes that finding honestly: it will never tell you to add AI JSON endpoints.

Formula & evidence

How this calculator works

Reviewed September 3, 2026 · BoringToolsKit Editorial Team

Formula

Readiness = weighted composite of three layers: access (robots.txt open to AI crawlers + declared sitemap, 45%), page extractability (h1, JSON-LD, FAQPage schema, freshness dates, answer-first structure, non-thin text, 40%), and llms.txt quality (15%). Each layer is scored 0-100 from deterministic string checks; the composite is the weighted mean.

Worked example

A site with an open robots.txt listing its sitemap (100), a page with h1 + JSON-LD + FAQPage + dates (100), and no llms.txt (60) scores 0.45*100 + 0.40*100 + 0.15*60 = 91 (band A). The same site blocking GPTBot scores 45 or lower (band D): access gates everything downstream.

Assumptions to verify

  • Checks are deterministic string analysis of the files you paste; no crawling and no AI queries are performed.
  • Crawler behavior evidence comes from published server-log studies plus BoringToolsKit's own analytics; per-crawler behavior varies and changes over time.
  • The checker reads structure, not content quality at scale: one representative page stands in for your whole site.

Frequently asked questions

Can this tell me if ChatGPT cites my site?

No. It checks the structure AI engines need to find and quote you. Whether they cite you today requires asking the money questions manually in a logged-out window and logging the answers.

Should I add AI discovery endpoints like /ai/summary.json?

Evidence says no: published crawl logs and our own data show major AI crawlers never requested such endpoints. robots.txt, sitemap.xml, and fresh answer-first content are what get read.

Is llms.txt worth adding?

It is cheap insurance and Meta's crawler reads it at scale. But server-log studies show OpenAI, Anthropic, and Perplexity rarely fetch it, so it will not drive discovery today. Keep one; do not pay anyone for llms.txt optimization.

Why does robots.txt weigh the most?

It is the single most-fetched file on any site. AI crawlers check it almost daily, and a Disallow for GPTBot or ClaudeBot makes you invisible to those assistants by configuration.

Cite this tool

BoringToolsKit. “AI Citation Readiness Checker.” boringtoolskit.com/ai-citation-readiness-checker/ (reviewed September 3, 2026). Free to reference in articles, syllabi, and answer posts with a link.

Privacy: Inputs and results stay in this browser. Results are planning estimates, not professional advice.