Site Audit

What the crawl checks, how domain verification works and how scores are reported.


Site Audit fetches pages from your site and returns page-level findings with the evidence attached.

#Verify your domain

Anyone can point a crawler at a website, so Citeroot limits unverified domains to the homepage plus a handful of pages (six in total). To audit up to your plan's page limit, prove you control the domain with one of:

  1. DNS TXT record — add a TXT record at _citeroot-verify.your-domain.com with the value citeroot-verify=<your token>. DNS changes can take up to an hour.
  2. A file — publish your token as the entire contents of https://your-domain.com/.well-known/citeroot-verify.txt.
  3. A meta tag — add <meta name="citeroot-verify" content="<your token>"> to the <head> of your homepage.

The Site audit page shows your token and the exact values, with a Check now button that explains what it found for each method. Changing the project's domain resets verification.

#What's discovered and checked

Citeroot reads your robots.txt and sitemap (including sitemap indexes, compressed sitemaps excepted) to find pages. If the sitemap lists fewer pages than your limit it adds links found on the homepage. It stays on your host and prefers shallower pages. Per-audit page limits: 25 on Scout, 100 on Trail, 250 on Summit, 1,000 on Expedition.

GroupExamples
Indexingnoindex directives, canonical references, response codes, redirects
Structuretitles, descriptions, H1s, heading order, language
Contentmain text present in the raw HTML, opening answer, alt text, internal links
Markupstructured data (JSON-LD) types and syntax, Open Graph
Crawler accessrobots.txt for 20 AI user agents, by purpose, and llms.txt

Pages are read as raw HTML, as most retrieval systems do. Pages that depend entirely on client-side scripts will look empty — which is itself a finding.

#Scores and findings

Each page gets a score from its checks; the site score weights shallower pages more. Findings are grouped by check across pages, so the template issue affecting 40 pages appears once with a count. Each check shows pass, needs attention or fail, the evidence and a suggested fix. Comparing two audits shows whether the score moved.

#A polite crawler

The crawler identifies itself as CiterootBot, honours robots.txt for its own user agent, caps speed and page size, only fetches public addresses, and re-checks every redirect. If your site blocks it, the audit says so and tells you how to allow it — which is also a hint about whether AI crawlers can get in. See /bot.

#Try it free

The page-readiness audit runs the single-page checks on any URL with no account.

Last updated Oct 8, 2026 · Suggest an edit