Robots.txt is one of the oldest tools in SEO and one of the most reliably misunderstood. The gap between what most practitioners think it does and what Google actually does with it is wide enough to create real ranking damage — the kind that’s hard to trace because the affected pages look fine in Search Console.
The Core Confusion: Crawl Blocking Is Not Index Blocking
When you disallow a URL in robots.txt, you are telling Googlebot not to fetch the page. You are not telling Google to remove it from the index or stop it from ranking.
Google has been explicit about this since at least 2019. If a disallowed URL has external links pointing to it, Google can index it on the basis of those links alone — without ever reading the page. The result: a URL in the index with no title, no snippet, sometimes just the raw URL and a description pulled from anchor text elsewhere. It will occasionally rank for branded queries. You’ll see impressions in Search Console. The page is a ghost — present in Google’s index, invisible to Googlebot, entirely outside your control.
We’ve run into this during keyword cannibalization work between two competing properties. URLs that were supposed to be suppressed via robots.txt were still showing up as index candidates, competing against the canonical pages we actually wanted to rank. The disallow had blocked crawling perfectly. The indexing had happened anyway.
How Google Actually Parses the File
Google follows RFC 9309, the formal specification for the robots exclusion protocol standardized in 2022. A few parsing behaviors most SEOs don’t account for:
- Longest match wins, not first match. If you have
Disallow: /blog/andAllow: /blog/featured/, Google applies the most specific rule. The order of rules in the file doesn’t determine which one wins — specificity does. - Googlebot ignores the wildcard block if a Googlebot-specific block exists anywhere in the file. A
DisallowunderUser-agent: *only applies to Googlebot when there’s no separateUser-agent: Googlebotsection. If that section exists, the wildcard is ignored entirely for Google’s crawlers. This breaks a lot of assumptions baked into templated robots.txt implementations. - The file is cached for up to 24 hours. Changes you make today may not affect Googlebot’s behavior until tomorrow. If you’re trying to block a page urgently — leaked staging content, say — robots.txt is a slow lever.
- Google caps the file at 500 kibibytes. Rules past that limit are silently ignored. If your file has accumulated years of ad-hoc additions from multiple teams, check the actual file size.
What Robots.txt Is Actually Useful For
The right use case is crawl efficiency, not content suppression. Blocking Googlebot from faceted navigation parameters, internal search results pages, or session-ID URLs preserves crawl budget for pages you actually want indexed.
On a large e-commerce site with aggressive faceting — color, size, price range, rating — the URL space can expand into the millions. Most of those URLs have near-zero additional search value. If Googlebot is working through /category?color=blue&size=M&sort=price_asc, that’s time not spent on new product pages or updated content. Blocking the parameter combinations in robots.txt, combined with a canonical strategy, directly shifts crawl attention toward pages that matter.
If you’re using robots.txt to suppress a staging environment or hide a page from searchers, you’re using the wrong tool. Use noindex in the page’s HTTP header or meta robots tag. Password protection is better still for staging. Robots.txt stops crawling. It does not stop indexing.
The Noindex-Plus-Disallow Trap
Here’s a failure mode that catches practitioners who know the crawl/index distinction but try to solve it wrong.
You add noindex to a page. You also add it to robots.txt as a precaution. The problem: Googlebot can’t read the noindex directive if it can’t crawl the page. The noindex tag is rendered HTML — it requires a fetch. So you’ve created a situation where the page can’t be de-indexed through the meta tag signal, because you’ve blocked the only mechanism Google has for reading it.
Google has confirmed this directly. If you want a page out of the index, you need to allow crawling so Googlebot can read the noindex tag. The disallow actively prevents the removal you’re trying to accomplish. Remove the disallow, let Googlebot crawl, wait for the page to drop from the index, and only then consider adding the disallow back — though at that point you usually don’t need it.
Wildcards and Path Matching: Where the Syntax Gets Slippery
Google supports * and $ as pattern-matching characters. Most other crawlers have incomplete or inconsistent support for both. If you’re using wildcard patterns to block URL parameter strings, what you see in a third-party robots.txt tester may not match what Googlebot actually does.
Disallow: /*?* blocks all parameterized URLs for Googlebot. It may not work for Bingbot, which follows a different extension of the robots exclusion standard. If cross-engine behavior matters, test against each engine’s documentation rather than assuming one ruleset covers all.
Disallow: /*.pdf$ blocks PDFs specifically. The $ anchors the match to the end of the URL. Without it, Disallow: /*.pdf also blocks /pdf-guide/, which isn’t a PDF. Small syntax difference, large unintended consequence.
Auditing Robots.txt as an SEO Signal, Not Just a Config File
Most robots.txt audits check for one thing: is important content accidentally disallowed? Necessary, but not sufficient. A complete audit looks at three things:
- Are there disallowed URLs that have external backlinks? If yes, those URLs may already be in Google’s index despite the disallow.
- Does the file have a Googlebot-specific block that overrides the wildcard block in ways you didn’t intend?
- Is any
noindexcontent also disallowed — making the noindex unreadable by Googlebot?
Cross-reference disallowed paths against your backlink profile — Ahrefs and Semrush both surface this — to find URLs that are disallowed but linked. Run those against a site: search to confirm which are actually indexed. The overlap is usually larger than expected.
Google Search Console’s URL Inspection tool will tell you directly whether a given URL is blocked by robots.txt and whether it’s indexed. Those two signals together tell you which failure mode you’re dealing with.
One Caveat Worth Sitting With
If your robots.txt strategy was built primarily to suppress content from search results rather than manage crawl access, audit it with that distinction front of mind. The disallow directive is doing exactly what it was designed to do — controlling crawl, not index presence. Those are different problems, and only one of them robots.txt can solve.

