Crawler access
How Citeroot reports AI crawler permissions by purpose and what the statuses mean.
Site Audit tests your robots.txt against 20 AI-related user agents, grouped by what each is for.
#Purposes
- Search & answer indexing — crawlers that build indexes assistants retrieve from.
- User-triggered fetches — requests made because a person asked an assistant about a page.
- Training & model improvement — crawlers and robots tokens that control training use.
#Statuses
- Allowed — the tested paths are allowed.
- Partly blocked — some tested paths are disallowed.
- Blocked — all tested paths are disallowed.
If robots.txt can't be fetched at all (for example your server returns an error for it), the whole report says it couldn't be assessed rather than guessing. That is never shown as "allowed".
Each result shows the rule that matched, written as Allow: /path or Disallow: /path, and the user-agent group it came from.
#Matching rules
Evaluation follows RFC 9309: the most specific user-agent group applies; within a group the longest matching path wins; Allow beats Disallow when equally specific; * and $ wildcards are supported.
#Beyond robots.txt
robots.txt isn't the only gate. CDN or WAF rules, rate limits, login walls and client-side rendering can all prevent access, and Citeroot only sees what its own crawler is served. If CiterootBot receives a 401, 403 or 429, the audit reports it as a possible block.
#Free checker
Try the AI crawler checker without an account, and read the crawler access guide.
Last updated Oct 8, 2026 · Suggest an edit