Two mechanisms, two failure modes
Indexing and training crawlers — GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot and others — visit on their own schedule, at scale, and build the corpus a model draws on. Being excluded here means long-term absence from a system's background knowledge of your market.
Live retrieval agents — the fetchers behind ChatGPT search, Perplexity, and browsing-enabled assistants — request your page in real time while composing an answer for a specific user. Being excluded here means absence from the exact moment a buyer is asking for a shortlist.
Most sites fail one and not the other, which is why a single "are we indexed?" check is not a diagnosis.
What a crawler actually does
- Resolves DNS and requests the URL with its own user agent string.
- Passes — or does not pass — your CDN, WAF, rate limiter, and bot-management rules.
- Reads robots.txt and applies any directive addressed to it.
- Receives HTML. Some execute JavaScript; many do not, or do so with strict time budgets.
- Extracts text, headings, links, and structured data from whatever it received.
- Follows internal links and sitemap entries to discover further pages.
Every step is a place to lose. The most common losses happen at steps two and four, and neither shows up in human analytics.
The quiet blockers
- Bot management defaults. Managed challenge rules on a CDN can serve a JavaScript interstitial to any unfamiliar agent. The crawler records an empty page, not a block.
- Aggressive rate limiting. A crawler that gets 429s on its first few requests may not return for weeks.
- Client-side rendering. If content only exists after hydration, a non-rendering fetcher sees an empty shell. Prerendered or server-rendered HTML solves this permanently.
- Geo or ASN blocking. Datacenter IP ranges are routinely restricted; crawlers live in them.
- Login and cookie walls. Anything gated is functionally invisible.
- Redirect chains and orphaned pages. Your best proof page is unreachable if nothing links to it and the sitemap omits it.
How to verify access rather than assume it
- Request your key pages with an AI crawler user agent and compare the response to what a browser sees.
- Check whether substantive text is present in the raw HTML with JavaScript disabled.
- Review server and CDN logs for GPTBot, ClaudeBot, PerplexityBot, and Google-Extended hits — presence is the proof, robots.txt is only the permission.
- Confirm your sitemap is declared in robots.txt and contains every page you want cited.
- Run the free AI Machine Readability Checker, which flags blocking issues before anything else.
Ongoing observation of which crawlers actually arrive is what BotIQ™ — AI crawler accessibility and activity monitoring — was built for. Directives declare intent; logs record reality.
Access is necessary and insufficient
Once crawlers reliably reach and render your pages, the next constraint is interpretation, not access. That is covered in Being Crawlable Is Not the Same as Being Understood. For the permissions layer specifically, see Why robots.txt Is Only Part of the AI Readability Problem, and for the full sequence, the pillar: Can AI Actually Read Your Website?
Run the AI Machine Readability Checker
Get a 100-point score across AI access, content structure, entity clarity, structured signals and authority evidence — with the exact gaps holding your pages back.
Check your website freeMore in the AI Machine Readability cluster
Can AI Actually Read Your Website?
The pillar guide — the five questions hiding inside "can AI read my site", in order.
Being Crawlable Is Not the Same as Being Understood
Access is the floor, not the finish line. What separates a fetched page from an interpreted one.
Why robots.txt Is Only Part of the AI Readability Problem
What a permissive robots.txt answers, and the larger question it leaves open.
How Structured Data Helps AI Understand Your Business
Schema labels information that already exists. What it can do, and what it cannot rescue.
What Makes a Website Machine-Readable for AI Search
The practical checklist across all five categories, with common failure modes for each.
Decision & Trust Cluster
Related Buyer & Recommendation Guides
High-intent reading on choosing an AI visibility partner, how recommendation confidence is built, and where real-world AI gaps show up.
- → Why Companies Choose Monic AI Systems
Recommendation reasoning, evidence architecture, and the buyer-intent signals that make a business defensibly recommendable by AI.
- → Questions to Ask Before Hiring a GEO Consultant
Due-diligence framework for evaluating AI visibility and GEO providers — recommendation visibility, corroboration, and implementation evidence.
- → When AI Visibility Consulting Is Not Needed
A candid look at the foundational digital maturity that has to be in place before AI visibility work pays back.
- → Monic AI Systems vs Traditional Search Agencies
Rankings vs recommendation systems, keyword optimization vs evidence architecture, static content vs expertise architecture.
- → Real AI Visibility Gaps We Uncovered
Recurring proof series: what AI understood, what it missed, where recommendation failures and abstention occurred — and how we closed the gap.
- → What AI Could Not Answer Before Founder Interviews
How founder-level expertise surfaces the reasoning AI systems abstain on — and how that reshapes recommendation quality.
- → AI Recommendation Confidence Framework
The Seen → Recommended → Chosen framework: how AI trust signals, corroboration, and abstention determine who gets selected.
Get discovered by AI.
Join The Weekly Firehose for weekly AI visibility insights, research, and practical strategies to help your business become the answer AI recommends.