Only missing parts are filled in: the scheme defaults to https and the path is always replaced by robots.txt at the root.
Fetch route: by default the request goes through this site's own same-origin forwarder (nginx turns it into nothing but https://target/robots.txt); only if that fails does it try a direct browser request and then fall back to public third-party relays. The result always says which route served it.
Per-route attempt log
Groups
Consecutive User-agent lines share the rules below them; each group lists its own Allow / Disallow / Crawl-delay and other directives.
Sitemap registry
Sitemap lines are not tied to a group; every one declared anywhere in the file is collected here.
This robots.txt declares no Sitemap.
Path permission test
Give a path and a crawler name: the verdict follows the standard rules (longest matching pattern wins, Allow breaks ties, a bot without its own group falls back to *), and the line below names the rule that decided it.
No result yet - check a site above first, the test needs its rules.
Syntax and SEO audit
Checked against how crawlers really behave: a missing leading slash, an empty Disallow, directives before any User-agent, unencoded non-ASCII paths, lines over 500 characters, files over 500KB, duplicate or conflicting rules, and noindex, which no crawler executes.
Raw text
What the status code means
The status code of robots.txt is part of the protocol: Google acts on it as below, and a 5xx even pauses crawling of the whole site.
| Status | What crawlers do |
|---|---|
200 | The file is read normally; a 200 with an empty body means "no restrictions at all" - everything is crawlable. |
404 | No robots.txt: the whole site is treated as crawlable, and crawlers cache that answer instead of probing again every time. |
401 / 403 | The rules cannot be read. Google carries on as if there were no file, while Bing and some others treat it as disallow-all. To really keep crawlers out, use authentication or a firewall rather than a 403 robots.txt. |
5xx | Server error: Google pauses crawling of the entire site for up to about an hour, because it cannot tell which paths are disallowed. A healthy site whose robots.txt returns 500 is the most easily missed incident. |
301 / 302 | The crawler follows the redirect and reads the file there; when it redirects to a different host, that host's rules are the ones in effect (this tool shows you the final status too). |
What is this
RobotsCheck fetches the /robots.txt at the root of a site you name and parses it per RFC 9309 into "user-agent groups + rules": each group lists its Disallow / Allow / Crawl-delay and other directives, unknown ones kept verbatim, and next to it you get a path permission test, a syntax audit and the raw text with line numbers. The most useful part is the tester: give a path and a crawler name and it applies the standard algorithm - longest matching pattern wins, Allow breaks ties, a bot with no group of its own falls back to * - then tells you which single line decided it.
Parsing, testing and auditing all happen in your browser. Fetching goes through one restricted same-origin forwarder on this site: nginx will only ever turn the request into https://<target>/robots.txt, refuses ports and private addresses, and has its access log switched off; nothing is cached. Deep paths, query strings and any user:pass@ credentials are dropped during normalisation and never enter the request. So this is not a fully offline tool, but it does not need any third-party service to work.
Features
- URL normalisationexample.com, https://example.com/deep/path?x=1 and even user:pass@ forms all normalise to https://example.com/robots.txt; only http and https schemes are accepted.
- RFC 9309 group parsingBOM, CRLF and # comments are stripped, field names are case-insensitive, consecutive User-agent lines merge into one group, and Allow / Disallow / Crawl-delay / Sitemap plus unknown directives are sorted out with the raw text preserved.
- Path permission testerSupports the * wildcard and the $ end anchor, picks the longest matching rule, lets Allow break ties, reads an empty Disallow: as "this group allows everything" and a missing group as allowed - then explains which rule matched and why.
- Syntax and SEO auditMissing leading slash, directives before any User-agent, unencoded non-ASCII paths, lines longer than 500 characters that crawlers truncate, files over 500KB that Google only reads partially, duplicate or conflicting rules, and noindex, which no crawler supports.
- Honest fetch routingThis site's forwarder first, then a direct attempt and finally public relays, with the serving route, HTTP status, byte size, line count and elapsed time listed - plus an expandable per-route attempt log, so you can tell a target-site problem from a route problem.
- Line-numbered raw textThe raw file is shown with line numbers and light highlighting of directive names, values and comments; everything fetched is escaped before it touches the page. Copy or download it as .txt, in Simplified Chinese, Traditional Chinese or English.
How to use
- Type a domain or any URL (pasting a full link is fine) and press Check robots.txt - or click one of the quick samples.
- Read the summary and the stat cards first: which route served it, the HTTP status, the number of groups, rules and sitemaps, and the overall verdict; the attempt log under it tells you whether the direct fetch was blocked.
- In the Path permission test, enter a path you care about and a crawler name, e.g. /wp-admin/ and Googlebot, and look at the verdict plus the rule that produced it.
- Fix the items in the syntax audit one by one, then use Copy raw / Download .txt to keep the before-and-after text for comparison.
Data source and scope
Parsing and matching follow RFC 9309, the robots.txt specification Google contributed to the IETF in 2022: field names are case-insensitive, the longest pattern by line length wins, Allow breaks ties and an empty Disallow means no restriction. The * wildcard and the $ end anchor are long-standing Google extensions supported by most major crawlers and implemented here too. Crawl-delay is honoured only by some crawlers (Bing, Yandex, Seznam) and Google explicitly ignores it; Host is a Yandex / Baidu extension; noindex has never been a robots.txt directive and nothing executes it.
The analysed content is the site's own published robots.txt; nothing is cached and the forwarding nginx location logs no requests. The fallback relay list lives in one constant, RELAYS, at the top of RobotsCheck/app.js?v=e18bb54c, so maintaining it means editing that one place. The audit thresholds are the documented numbers: Google reads only the first 500KB of robots.txt, cuts lines past 500 characters, and requires non-ASCII characters to be percent-encoded.
FAQ
- Why not let the browser fetch it directly?
- The same-origin policy stops a page from reading a response from another origin unless that server sends CORS headers, and almost no site puts them on robots.txt. So the tool asks this site to forward the request once (nothing but https://<domain>/robots.txt, ports and private addresses refused), and the summary states which route answered.
- Does forwarding through this site expose what I checked?
- The forwarding nginx rule has its access log switched off, nothing is cached or stored, and the domain you asked about exists on the server only for the moment that one request is in flight. Deep paths, query strings and credentials are dropped during normalisation and never leave your browser. Only the fallback relays put the request on a third-party server, and those are not used unless the forwarder fails.
- Does the verdict match what Google decides?
- On the core algorithm yes: longest match wins, Allow breaks ties, * is a wildcard, $ anchors the end, and a bot with no group of its own falls back to *. Google additionally normalises URLs before comparing (case, percent-encoding equivalence, treating index.html and the directory form as the same URL) and handles 401/403 and 5xx specially; those cases are called out in the audit and the status table.
- Does disallowing a path in robots.txt actually keep crawlers out?
- Only crawlers that obey it. robots.txt is a request for permission, not access control: rule-breaking crawlers ignore it entirely, and since the file is public it hands outsiders your list of private paths. Use authentication or server-side rules to really restrict access. Also do not expect Disallow to remove an indexed page - a blocked URL can stay in the index as a result without a snippet.
- Why is the status code worth looking at?
- Because it is part of the protocol: 404 or an empty file means full access and crawlers cache that judgement; on 401/403 Google keeps crawling while some other crawlers assume disallow-all; and a 5xx makes Google pause crawling the whole site for up to an hour - a site whose pages are fine but whose robots.txt returns 500 is a classic invisible outage. Read the status together with the serving route.
More tools on this site
- IP Lookup (geo / ISP / ASN)
- Chinese Converter
- Text Diff & Review
- Venn Diagram Generator
- Image Compressor
- PDF Toolkit
Parsing and testing run in your browser · free · EN / 简体 / 繁體