NEW CASE
Antana

We create digital solutions that work for businesses


Give us a call +38 (066) 35-14-529

Let's take the first step towards your website — write to us

Close
Tools

Robots.txt Tester

Whether a crawler may fetch a specific URL, and which rule decided it.

What this tester does

It fetches your robots.txt, parses the rules and shows whether a crawler may reach a given URL. Crucially, it shows which rule decided.

A verdict for one specific URL Not “the file looks fine” but an answer for the page you care about. That is where surprises turn up.
The deciding rule We show the line and its number in the file. When the result surprises you, the cause is immediately visible.
Different crawlers Googlebot and GPTBot can get different answers for the same URL when each has its own rule group.
Syntax errors Lines without a colon, rules before the first User-agent, unknown fields — each renders part of the file inert.

How rule selection works

The logic is not obvious, and every misunderstanding starts here.

A specific group wins If a crawler has its own User-agent block, the * rules do not apply to it at all. They are replaced, not combined.
The longest match wins Among matching rules the one with the longer path applies. That is why Allow: /private/public/ overrides Disallow: /private/.
Ties go to Allow When lengths are equal, Allow beats Disallow. The specification says so explicitly.
The * and $ characters An asterisk matches any sequence, a dollar marks the end of the URL. Disallow: /*.pdf$ blocks the files but not URLs with parameters after them.
An empty Disallow A Disallow line with no value means “everything is allowed”. Not an error but the conventional way to open a site fully.

What robots.txt does not do

It does not hide a page from search Disallow forbids visiting, not displaying. If links point to the page, the URL can still surface without a description. Hiding requires noindex.
It does not protect data The file is public, so a list of forbidden sections is effectively a map of where to look. Private areas need a password.
It is not binding on everyone It is a convention. Search engines respect it; content scrapers often do not.
It does not cover subdomains Every subdomain has its own robots.txt. The main domain file does not apply to shop.example.com.

What to do with a surprising result

1 Look at the rule and line number Often the path is broader than you thought: Disallow: /cat blocks /catalog and /category too.
2 Check for a specific group If the crawler has its own User-agent block, the shared rules are ignored. You have to copy them across by hand.
3 Add an exception with Allow Rather than rewriting the ban, open the subsection with its own line. The longer rule wins.
4 Re-check after editing Edits take effect immediately, but Google re-reads the file with a delay — usually within a day.
Telegram Viber Call us