How to verify search and AI crawler access without exposing private pages
Check robots rules, indexing directives and hosting responses separately, then confirm the intended public pages are accessible while private routes stay protected.
A crawler access problem is easier to fix when you identify the exact layer responsible. A permissive robots file will not remove a hosting challenge, and a successful browser visit will not prove that a search crawler receives the same document. Treat access as a set of observable responses, with public and private routes considered separately.
Start with the intended access policy
List a few routes that should be public and a few that should remain private. An article, service page and contact page are sensible public examples. Account screens, payment records, unpublished previews and customer reports are different. For each route, state whether it should be retrievable, indexable and discoverable through internal links.
Do not use robots.txt to protect confidential information. Google describes it as a crawling control, and a blocked URL can still appear in results without its content. Authentication belongs at the application or hosting layer. Google’s robots.txt introduction explains this distinction.
Distinguish crawler purposes
“AI bot” is too broad for a reliable rule. OpenAI identifies OAI-SearchBot for search, GPTBot for potential model training and ChatGPT-User for user-initiated retrieval. Search and training choices are independent; ChatGPT-User is not the automatic search crawler, and user-triggered requests may not follow robots rules in the same way. Check OpenAI’s current crawler documentation before changing those settings.
Anthropic similarly documents Claude-SearchBot, Claude-User and ClaudeBot for search, user-directed retrieval and potential training respectively. Use the exact current names in its official crawler guidance. A policy about training should be recorded as a policy choice, not mislabeled as proof that search access has failed.
Inspect the origin’s actual robots file
Retrieve /robots.txt from the same scheme and hostname as the page. Confirm it is a text response containing your rules, rather than an HTML error or challenge. Preserve a copy before editing. Read the complete file, including named agent groups and wildcard groups, instead of searching only for one “Disallow” line.
Evaluate the relevant agent against the actual public path. A restriction on /account/ says little about /articles/. Conversely, a broad restriction can hide an otherwise healthy public section. Avoid pasting a universal “Allow all” replacement: it can discard valid restrictions and change policies you did not intend to change.
Check the page response and indexability
Retrieve the public URL and record its status, redirect chain, content type, title and a short identifying passage. Inspect both response headers and HTML for indexing controls. Confirm the document is the intended page and the canonical points to its deliberate preferred destination.
Google needs to crawl a page to discover its noindex instruction; blocking crawling can prevent that instruction from being read. This matters when intentionally removing a page from search. See Google’s noindex documentation. Do not remove noindex from an account screen merely because an audit flags it.
Investigate the hosting layer with logs
If a permitted public page receives a denial or challenge, inspect the corresponding hosting security event. Match the URL, time, response and rule that acted. A local request with a changed User-Agent is a diagnostic comparison, not proof that it came from a genuine vendor crawler. Anyone can send a header claiming to be a bot.
Use the provider’s current verification guidance and available request logs to establish identity. Adjust the narrow rule responsible for the unintended block, preserve authentication and rate controls, and avoid disabling every security feature across the zone. Retest the same route and keep the before-and-after evidence.
Close the loop without claiming indexing
A useful access record includes the exact URL, agent or request class, policy decision, observed response and follow-up check. Where available, inspect the page through the search engine’s verified webmaster property. Keep “retrieved successfully,” “indexed” and “shown for a query” as separate states.
After a template or hosting update, repeat these checks using the publishing validation checklist. Then return to the broader readiness review. A successful access check removes one obstacle; it does not promise a search position or an answer citation.
Sources and further reading
- Google Search Central: robots.txt introduction
- Google Search Central: block indexing with noindex
- OpenAI: overview of crawlers
- Anthropic: crawler purposes and controls
Provider documentation can change. Check the current guidance before changing crawler or search settings.