URL Access Testing Tool
Evaluates specific paths against User-Agent groups using wildcard specificity and conflict matching algorithms.
Multi-Crawler Comparison
Compare path access rules side-by-side for Googlebot, Bingbot, Yandex, Baidu, and generic search bots.
Sitemap Discovery & Verification
Extracts declared sitemap locations and checks live HTTP status codes and XML content headers.
Fetching & Auditing Robots.txt…
Evaluating origin URL, HTTP response headers, User-Agent groups, Allow/Disallow rule specificity, and sitemap declarations.
Robots.txt Analysis Completed
Target Origin: https://example.com/robots.txt
Site-Wide Crawl Restriction Detected (`Disallow: /`)
A Disallow directive blocking the root path / is present. This rule instructs crawlers not to crawl any URL on this site.
HTTP Status
200
Robots.txt FoundUser-Agent Groups
0
Crawler GroupsDisallow Rules
0
Crawl Block DirectivesAllow Rules
0
Explicit Allow RulesSitemaps Declared
0
Declared XML SitemapsEvaluate Path URL Access Live
Multi-Crawler Access Comparison Matrix
| Sample Path | Googlebot | Bingbot | Googlebot-Image | DuckDuckBot | YandexBot | Baiduspider | * (All) |
|---|
| Declared Sitemap URL | Line # | HTTP Verification Status | Content-Type |
|---|---|---|---|
| No sitemap declarations found in robots.txt. | |||
Line-Level Syntax Diagnostics & Warnings
Critical SEO Guidance & Safety Rule
robots.txt controls crawler access; it is not a reliable mechanism for removing an already indexed URL from search results or securing private data. Blocked URLs can still be indexed via external backlinks and can still be accessed directly by web browsers.
What is a Robots.txt File?
A robots.txt file is a plain text file placed at the root directory of a website (e.g. https://example.com/robots.txt). It provides crawl instructions to search engine web crawlers (like Googlebot, Bingbot) detailing which pages or directories they are permitted or forbidden to crawl.
Auditing your site's robots.txt file ensures that search engines are not accidentally blocked from indexing key revenue-generating landing pages or CSS/JS design assets.
Robots.txt Best Practices for SEO
Recommended Practices
- Always declare your primary XML sitemap URL inside robots.txt.
- Disallow administrative folders (`/admin/`, `/wp-admin/`, `/cart/`).
- Ensure CSS and JavaScript assets are accessible to Googlebot.
- Keep file size under 500 KB to guarantee parser readability.
Pitfalls to Avoid
- Never declare `Disallow: /` on live production websites.
- Do not use robots.txt to hide sensitive user security data.
- Avoid conflicting Disallow and Allow wildcard rules.