Short answer: blocking a URL in robots.txt does not remove it from Google. If a sensitive AI-share page, transcript, or generated conversation can be discovered through links, Google may still index the URL without being able to see a hidden noindex directive. For SEO and SGO teams, the fix is to use authentication, remove public links, return the right status code, or allow crawling long enough for noindex to be seen.
Search Engine Journal reported that shared Claude chats appeared in Google even though the pages were blocked by robots.txt; the key lesson was not about Claude alone, but about a common indexing mistake: disallow is not noindex. That distinction matters more as companies publish AI-generated share pages, prompt libraries, chat transcripts, support answers, and agent-accessible documentation.
What happened
The reported issue involved publicly shareable AI chat URLs that could show up in Google. SEJ’s review found a noindex signal sitting behind a robots.txt block. That creates a trap: if Googlebot is blocked from crawling the page, it may never fetch the page or response header needed to see the noindex instruction.
Google’s robots documentation has long made this clear: robots.txt controls crawling, not indexing. A blocked URL can still be indexed if Google discovers it through links or other signals, often with limited information because the page itself was not crawled.
Why this matters for AI search visibility
AI search programs create more semi-public URLs than traditional websites used to have: shared chats, AI-generated summaries, internal knowledge-base previews, prompt outputs, support answer drafts, and experimental agent pages. Many teams try to control them with one line in robots.txt. That may reduce crawling, but it does not reliably prevent search visibility.
- Privacy risk: a sensitive URL can be discoverable even when its content is not fully crawled.
- Brand risk: AI-generated or user-generated pages can become the version of your brand that search engines encounter.
- Measurement risk: teams may assume content is excluded when it is only crawl-blocked.
- SGO risk: answer engines need clean, intentional source material. Accidental URLs create noisy evidence.
The practical rule: crawl control and index control are different
| Goal | Use this | Avoid relying on |
|---|---|---|
| Stop bots from fetching low-value public paths | robots.txt | robots.txt as privacy protection |
| Remove a page from Google’s index | noindex visible to crawlers, 404/410, or removal workflows | A hidden noindex on a blocked URL |
| Protect private AI chats or documents | Authentication, access control, private-by-default sharing | Public share links plus crawler directives |
| Keep AI systems focused on approved sources | Canonical, indexable, internally linked content hubs | Loose collections of public generated pages |
What SEO and SGO teams should check now
- Inventory public AI-share URLs. Search for shared chat pages, public transcript URLs, prompt result pages, and preview links.
- Test whether blocked URLs are still known to Google. Use URL Inspection where available and review indexed URL samples.
- Do not hide
noindexbehindDisallow. If a page must be removed, let Google crawl the directive until the URL drops, then block crawling only if needed. - Use 404 or 410 for share pages that should no longer exist. Do not leave thin placeholder pages live.
- Move sensitive AI outputs behind login. Search directives are not access control.
- Strengthen canonical source pages. Link AI systems and crawlers toward intentional explainers, product pages, documentation, and citation-ready resources.
For broader implementation, use the AI Search Optimization Checklist, the SGO playbook, and the GEO guide to separate intentional AI-search assets from URLs that should not be discovered.
Editorial takeaway
The Claude example is a reminder that AI-search optimization is not only about getting cited. It is also about controlling which pages deserve to become evidence. If a URL should shape answers, make it crawlable, useful, and internally supported. If it should not be visible, protect it properly instead of assuming robots.txt will keep it out of search.
