Guide

AI crawler access: robots.txt, llms.txt and rendering

Which AI crawlers exist, what each is for, how to configure robots.txt by purpose, and why raw HTML still matters.

Before an AI system can use your content, it has to be able to fetch it. This guide covers the access layer: who's knocking, what you can say about it and the common ways sites accidentally shut the door.

#Who's knocking

Several providers publish separate user agents for separate jobs. Check each provider's current documentation, because names and behaviour change.

PurposeExamples
Search and answer indexingOAI-SearchBot (OpenAI), PerplexityBot (Perplexity), Claude-SearchBot (Anthropic), Googlebot, Bingbot
TrainingGPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended token, Applebot-Extended, CCBot
User-triggered fetchesChatGPT-User (OpenAI), Perplexity-User, Claude-User

A few clarifications:

  • Google-Extended is a robots.txt token, not a separate crawler. Google documents that it doesn't affect inclusion in Google Search.
  • Allowing a search crawler doesn't opt you into training, and vice versa, for providers that separate them.
  • User-triggered fetchers behave differently from background crawlers. Providers describe how, or whether, robots.txt applies; read their docs before assuming.

#Configure by purpose

robots.txt works per user agent. Group the bots you treat the same:

# Allow search and user-triggered access
User-agent: OAI-SearchBot
User-agent: PerplexityBot
User-agent: Claude-SearchBot
User-agent: ChatGPT-User
User-agent: Perplexity-User
User-agent: Claude-User
Allow: /

# Opt out of model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /

# Everyone else
User-agent: *
Disallow: /admin/
Sitemap: https://example.com/sitemap.xml

Rules are matched per RFC 9309: the most specific user-agent group applies, and within a group the longest matching path wins.

Whether to restrict training is a business decision. Make it deliberately and document it.

#robots.txt is not access control

It's an honour system for well-behaved crawlers. For private content use authentication. For abusive traffic use your firewall — but note that blanket WAF rules ("block anything with 'bot' in the user agent") can lock out the search crawlers you want. Verify a crawler's identity using the provider's published IP ranges or verification method before allowing it through.

#Rendering matters

Many retrieval systems read the HTML your server returns and don't execute JavaScript. Check that:

  • The main content and key facts appear in the raw HTML.
  • Important pages don't depend on client-side navigation to be reachable.
  • Response codes are correct (200 for pages, 404 for missing ones, no soft-404s).

A quick test: curl the page and read the output. If the answer isn't there, a crawler may not see it either. The page-readiness audit automates that check.

#What about llms.txt?

llms.txt is a proposed convention: a Markdown file at /llms.txt summarising your site for language models. It's not a standard that all engines honour, and Google's guidance says special AI text files aren't required for its search experiences. If you maintain one, treat it as documentation. Generate one in a minute with our llms.txt generator.

#Common mistakes

  • Staging rules (Disallow: /) shipped to production.
  • Allowing bots in robots.txt but blocking them in a CDN or WAF.
  • Blocking by path patterns that accidentally catch documentation or pricing pages.
  • Redirect chains on robots.txt itself.
  • Assuming a successful crawl means inclusion in answers. Access is necessary, not sufficient.

#Check your site

Run the free AI crawler checker. It tests your robots.txt against 20 AI-related user agents grouped by purpose and shows the matching rule for each.

#Sources

See what AI says about you today.

Add your site and a few questions your buyers ask. Get your first baseline in an afternoon — 14-day free trial · no card.