Robots.txt Tester
Check whether a URL is blocked, and see exactly which rule decided it.
About the Robots.txt Tester
Precedence is by length, not by order. The longest matching path wins regardless of where it sits in the file, and Allow beats Disallow when both match at the same length. Reading the file top to bottom and taking the first match — which is what almost everyone assumes — gives the wrong answer whenever a broad Disallow appears above a narrow Allow, which is exactly the shape most real robots.txt files have.
A crawler uses one group and ignores every other, including the wildcard. This is the mistake that does the most damage: adding a Googlebot section means Googlebot stops reading the * section entirely, so every rule it still needs has to be repeated inside its own group. Sites regularly open up paths to Google that they believed were blocked, because the block only ever existed in the wildcard group.
Consecutive User-agent lines share one set of rules. Two agents listed together followed by one Disallow applies that rule to both — not, as it reads, only to the second.
Worth remembering what robots.txt is for. It stops crawling, not indexing: a URL that is linked from elsewhere can still appear in results, listed without a description, precisely because the crawler was forbidden from fetching it to find out more. To keep a page out of the index, let it be crawled and use a noindex tag — a blocked page is one whose noindex tag can never be read.
How it works
Paste your robots.txt and the path you want to check.
Pick a crawler — the answer differs by agent more than people expect.
The deciding rule is named, with its line number.
Frequently asked questions
- Why is my URL still blocked when I have an Allow rule for it?
- Precedence goes to the longest matching path, not the first or last rule in the file. If a Disallow matches more characters of your URL than the Allow does, the Disallow wins wherever it appears. Make the Allow path more specific than the Disallow it needs to beat.
- Does a Googlebot section replace the wildcard section?
- Yes, entirely. A crawler picks the single most specific group that names it and ignores all the others, so once a Googlebot group exists, Googlebot never reads the * group again. Every rule you still want applied has to be repeated inside it.
- Does robots.txt stop a page being indexed?
- No — it stops it being crawled, which is different. A blocked URL that is linked from elsewhere can still be indexed and shown without a description. To keep something out of the index, allow crawling and use a noindex meta tag, because a blocked page is one whose noindex can never be read.
- What do the * and $ wildcards do?
- An asterisk matches any run of characters, so /*.pdf matches any PDF at any depth. A dollar anchors the match to the end of the URL, so /*.pdf$ matches file.pdf but not file.pdf?download=1. Both are supported by the major crawlers, though not by the original standard.
- Is my robots.txt uploaded anywhere?
- No. Parsing and matching both run in your browser, and nothing is fetched — you paste the file rather than giving a URL, which also means you can test a version before publishing it.
Privacy
Everything happens locally. Your files are read by your own browser, processed on your device, and never uploaded — closing the tab is all it takes to erase them.
Related tools
Robots.txt Generator
Build a valid robots.txt with per-crawler allow and disallow rules.
UTM Builder
Build tagged campaign URLs that report correctly, with the mistakes flagged as you type.
Word Density Checker
Analyse keyword frequency and density, including two- and three-word phrases.
Meta Tag Generator
Generate title, description, Open Graph and Twitter card tags, with live previews.
Schema Markup Generator
Build valid JSON-LD structured data for articles, products, FAQs, events and more.
Hreflang Generator
Build hreflang tags for a multi-language site, with the invalid codes caught first.