Website and SEO

Crawler access: can search engines and AI assistants read your site?

Why ChatGPT, Claude and Perplexity take no URL submissions, which crawlers they use, and how DomainGuard checks robots.txt and your firewall for each one every day.

Updated · 2 min read

You can tell Bing a page changed with IndexNow. You can't do that with ChatGPT, Claude or Perplexity: none of them takes URL submissions. They find pages two ways, with their own crawlers and through the search engines they lean on (ChatGPT search reads a lot of Bing). So getting into AI answers comes down to two things: the search engines have your pages, and the crawlers can get in.

Plan: Starter and up, like the rest of SEO.

The crawlers we check

Purpose Crawlers What a block costs
Search engines Googlebot, Bingbot, Applebot, AhrefsBot Google, Bing (and DuckDuckGo, Copilot), Siri and Spotlight, Yep
AI search OAI-SearchBot, Claude-SearchBot, PerplexityBot Whether ChatGPT, Claude or Perplexity can show and cite the page
AI opening a page someone asked about ChatGPT-User, Claude-User, Perplexity-User Whether the assistant can read the page when a person points it there
AI model training GPTBot, ClaudeBot, Google-Extended, Applebot-Extended Only training. Blocking these does not hide you from search or AI answers

What the check does

For each crawler we read your robots.txt the way the crawler does (RFC 9309: its own group first, then *, the longest matching rule wins). Then we load your home page with that crawler's user agent and compare the answer with what a normal browser gets. A firewall rule, bot fight mode or a "block AI bots" switch shows up there as a 403 or a challenge page.

The request comes from our address, not the crawler's. A firewall that trusts only a crawler's real addresses can treat it differently, so "OK" is a good sign, not proof. "Blocked" is a strong one. If your site didn't answer a browser either, we say so and don't count it against any crawler.

When it runs

Every domain with IndexNow on is checked once a day. If a search or AI crawler could get in at the last check and can't now, it goes in your morning summary. You can also check any time from the site's SEO page (Crawler access, under IndexNow) or ask your AI assistant: the MCP tools are get_crawler_access, check_crawler_access and list_crawler_access_history, and get_search_visibility_overview shows every domain's IndexNow setup and crawler access at once.

Fixing a block

  • robots.txt: remove the Disallow: / from the crawler's group, or from User-agent: * if it has no group of its own. Disallowing private paths like /admin is fine.
  • Cloudflare: under Security → Settings, check the AI bot settings, AI Labyrinth and Bot fight mode. Managed robots.txt can add AI blocks to your file without you editing it.
  • Run Check now afterwards to confirm.

Still stuck?

Ask the people who run the scanner.

Send the domain and what you expected to see. We look at the same scan you are looking at and write back with what it means and what to change.