← Back to Research Nav

robots.txt Checker

Group parsing · path verdict · syntax audit

English 简体 繁體

Only missing parts are filled in: the scheme defaults to https and the path is always replaced by robots.txt at the root.

Quick samples

Fetch route: by default the request goes through this site's own same-origin forwarder (nginx turns it into nothing but https://target/robots.txt); only if that fails does it try a direct browser request and then fall back to public third-party relays. The result always says which route served it.

Path permission test

Give a path and a crawler name: the verdict follows the standard rules (longest matching pattern wins, Allow breaks ties, a bot without its own group falls back to *), and the line below names the rule that decided it.

No result yet - check a site above first, the test needs its rules.

What the status code means

The status code of robots.txt is part of the protocol: Google acts on it as below, and a 5xx even pauses crawling of the whole site.

Status What crawlers do
200The file is read normally; a 200 with an empty body means "no restrictions at all" - everything is crawlable.
404No robots.txt: the whole site is treated as crawlable, and crawlers cache that answer instead of probing again every time.
401 / 403The rules cannot be read. Google carries on as if there were no file, while Bing and some others treat it as disallow-all. To really keep crawlers out, use authentication or a firewall rather than a 403 robots.txt.
5xxServer error: Google pauses crawling of the entire site for up to about an hour, because it cannot tell which paths are disallowed. A healthy site whose robots.txt returns 500 is the most easily missed incident.
301 / 302The crawler follows the redirect and reads the file there; when it redirects to a different host, that host's rules are the ones in effect (this tool shows you the final status too).

What is this

RobotsCheck fetches the /robots.txt at the root of a site you name and parses it per RFC 9309 into "user-agent groups + rules": each group lists its Disallow / Allow / Crawl-delay and other directives, unknown ones kept verbatim, and next to it you get a path permission test, a syntax audit and the raw text with line numbers. The most useful part is the tester: give a path and a crawler name and it applies the standard algorithm - longest matching pattern wins, Allow breaks ties, a bot with no group of its own falls back to * - then tells you which single line decided it.

Parsing, testing and auditing all happen in your browser. Fetching goes through one restricted same-origin forwarder on this site: nginx will only ever turn the request into https://<target>/robots.txt, refuses ports and private addresses, and has its access log switched off; nothing is cached. Deep paths, query strings and any user:pass@ credentials are dropped during normalisation and never enter the request. So this is not a fully offline tool, but it does not need any third-party service to work.

Features

How to use

  1. Type a domain or any URL (pasting a full link is fine) and press Check robots.txt - or click one of the quick samples.
  2. Read the summary and the stat cards first: which route served it, the HTTP status, the number of groups, rules and sitemaps, and the overall verdict; the attempt log under it tells you whether the direct fetch was blocked.
  3. In the Path permission test, enter a path you care about and a crawler name, e.g. /wp-admin/ and Googlebot, and look at the verdict plus the rule that produced it.
  4. Fix the items in the syntax audit one by one, then use Copy raw / Download .txt to keep the before-and-after text for comparison.

Data source and scope

Parsing and matching follow RFC 9309, the robots.txt specification Google contributed to the IETF in 2022: field names are case-insensitive, the longest pattern by line length wins, Allow breaks ties and an empty Disallow means no restriction. The * wildcard and the $ end anchor are long-standing Google extensions supported by most major crawlers and implemented here too. Crawl-delay is honoured only by some crawlers (Bing, Yandex, Seznam) and Google explicitly ignores it; Host is a Yandex / Baidu extension; noindex has never been a robots.txt directive and nothing executes it.

The analysed content is the site's own published robots.txt; nothing is cached and the forwarding nginx location logs no requests. The fallback relay list lives in one constant, RELAYS, at the top of RobotsCheck/app.js?v=e18bb54c, so maintaining it means editing that one place. The audit thresholds are the documented numbers: Google reads only the first 500KB of robots.txt, cuts lines past 500 characters, and requires non-ASCII characters to be percent-encoded.

FAQ

Why not let the browser fetch it directly?
The same-origin policy stops a page from reading a response from another origin unless that server sends CORS headers, and almost no site puts them on robots.txt. So the tool asks this site to forward the request once (nothing but https://<domain>/robots.txt, ports and private addresses refused), and the summary states which route answered.
Does forwarding through this site expose what I checked?
The forwarding nginx rule has its access log switched off, nothing is cached or stored, and the domain you asked about exists on the server only for the moment that one request is in flight. Deep paths, query strings and credentials are dropped during normalisation and never leave your browser. Only the fallback relays put the request on a third-party server, and those are not used unless the forwarder fails.
Does the verdict match what Google decides?
On the core algorithm yes: longest match wins, Allow breaks ties, * is a wildcard, $ anchors the end, and a bot with no group of its own falls back to *. Google additionally normalises URLs before comparing (case, percent-encoding equivalence, treating index.html and the directory form as the same URL) and handles 401/403 and 5xx specially; those cases are called out in the audit and the status table.
Does disallowing a path in robots.txt actually keep crawlers out?
Only crawlers that obey it. robots.txt is a request for permission, not access control: rule-breaking crawlers ignore it entirely, and since the file is public it hands outsiders your list of private paths. Use authentication or server-side rules to really restrict access. Also do not expect Disallow to remove an indexed page - a blocked URL can stay in the index as a result without a snippet.
Why is the status code worth looking at?
Because it is part of the protocol: 404 or an empty file means full access and crawlers cache that judgement; on 401/403 Google keeps crawling while some other crawlers assume disallow-all; and a 5xx makes Google pause crawling the whole site for up to an hour - a site whose pages are fine but whose robots.txt returns 500 is a classic invisible outage. Read the status together with the serving route.

More tools on this site

Parsing and testing run in your browser · free · EN / 简体 / 繁體