Website and SEO
Crawler access: can search engines and AI assistants read your site?
Why ChatGPT, Claude and Perplexity take no URL submissions, which crawlers they use, and how DomainGuard checks robots.txt and your firewall for each one every day.
Updated · 2 min read
You can tell Bing a page changed with IndexNow. You can't do that with ChatGPT, Claude or Perplexity: none of them takes URL submissions. They find pages two ways, with their own crawlers and through the search engines they lean on (ChatGPT search reads a lot of Bing). So getting into AI answers comes down to two things: the search engines have your pages, and the crawlers can get in.
Plan: Starter and up, like the rest of SEO.
The crawlers we check
| Purpose | Crawlers | What a block costs |
|---|---|---|
| Search engines | Googlebot, Bingbot, Applebot, AhrefsBot | Google, Bing (and DuckDuckGo, Copilot), Siri and Spotlight, Yep |
| AI search | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Whether ChatGPT, Claude or Perplexity can show and cite the page |
| AI opening a page someone asked about | ChatGPT-User, Claude-User, Perplexity-User | Whether the assistant can read the page when a person points it there |
| AI model training | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended | Only training. Blocking these does not hide you from search or AI answers |
What the check does
For each crawler we read your robots.txt the way the crawler does (RFC 9309: its own group first, then *, the longest matching rule wins). Then we load your home page with that crawler's user agent and compare the answer with what a normal browser gets. A firewall rule, bot fight mode or a "block AI bots" switch shows up there as a 403 or a challenge page.
The request comes from our address, not the crawler's. A firewall that trusts only a crawler's real addresses can treat it differently, so "OK" is a good sign, not proof. "Blocked" is a strong one. If your site didn't answer a browser either, we say so and don't count it against any crawler.
When it runs
Every domain with IndexNow on is checked once a day. If a search or AI crawler could get in at the last check and can't now, it goes in your morning summary. You can also check any time from the site's SEO page (Crawler access, under IndexNow) or ask your AI assistant: the MCP tools are get_crawler_access, check_crawler_access and list_crawler_access_history, and get_search_visibility_overview shows every domain's IndexNow setup and crawler access at once.
Fixing a block
- robots.txt: remove the
Disallow: /from the crawler's group, or fromUser-agent: *if it has no group of its own. Disallowing private paths like/adminis fine. - Cloudflare: under Security → Settings, check the AI bot settings, AI Labyrinth and Bot fight mode. Managed robots.txt can add AI blocks to your file without you editing it.
- Run Check now afterwards to confirm.
Keep reading
Related articles
- IndexNow: tell search engines a page changedTurning IndexNow on for a domain, the key file your site has to serve, which engines get notified, sending pages by hand or from the sitemap, and the record of every URL sent and what each engine answered.Website and SEO ·Updated
- MCP tool referenceAll 237 tools a customer key can be offered, by scope, with each one's access level. 128 read, 100 write, 9 destructive; 7 spend credits.API and MCP ·Updated
- Search intent reportEvery crawled page read as informational, commercial, transactional or navigational; pairs competing for one search with a fix; missing links; SERP.Website and SEO ·Updated
