Robots.txt Checker

Fetch and validate a robots.txt and test whether a URL is blocked.

What does a robots.txt checker do?

A robots.txt checker downloads a site's robots.txt, finds the rules that apply to a chosen crawler such as Googlebot, and tells you whether a specific URL path is allowed or blocked — and which rule decided it.

How rules are matched

  1. Pick the group: the crawler uses the group with the most specific matching User-agent (e.g. Googlebot) and falls back to User-agent: *.
  2. Find matching rules: every Allow and Disallow whose pattern matches the start of the path, with * as a wildcard and $ as end-of-URL.
  3. Longest rule wins. If Disallow: /shop and Allow: /shop/sale both match /shop/sale/shoes, the longer Allow wins. On a tie, Allow wins.
  4. No matching rule means the path is allowed.

This checker follows the same logic Google documents in RFC 9309.

Common robots.txt mistakes

  • Disallow: / left over from staging — blocks the entire site. This is the first thing to check after a launch.
  • Blocking CSS/JS folders such as /wp-includes/ or /assets/, which stops Google rendering pages properly.
  • Using robots.txt to remove pages from Google — it does not de-index; use noindex.
  • Case and trailing slashes — paths are case-sensitive: Disallow: /Admin does not block /admin.
  • Missing Sitemap line — not required, but it helps all crawlers find your sitemap.

Reading the result

You see whether robots.txt exists, whether a sitemap is declared, whether the site root is crawlable, the verdict for your test path with the deciding rule, every rule applied to the chosen crawler, and the raw file. If the file returns 404, crawlers treat the whole site as allowed.

Need a new file? Use the Robots.txt Generator. Then validate the sitemap with the Sitemap Checker.

Frequently asked questions

How do I check if a URL is blocked by robots.txt?

Enter the website, type the path to test (for example /blog/post) and pick a crawler. The checker shows Allowed or Blocked and the rule that decided it.

What happens if robots.txt does not exist?

Crawlers assume everything is allowed. A missing file is not an error, but adding one with a Sitemap line is good practice.

Which rule wins when Allow and Disallow both match?

The most specific (longest) rule wins. If they are equally long, Allow wins.

Is robots.txt case-sensitive?

Paths are case-sensitive, so Disallow: /Admin does not block /admin. User-agent names are matched case-insensitively.

Why does Google still index a blocked page?

robots.txt blocks crawling, not indexing. If other sites link to the URL, Google may index it without content. Use noindex on a crawlable page instead.