Crawlability Checker

Audit search engine bot crawlability, robots.txt Disallow rules, meta robots noindex directives, HTTP status codes, and redirect chains.

Popular examples: https://google.comhttps://github.com
Crawlability vs Indexability

Separates bot fetch permission (robots.txt) from indexing permission (meta noindex).

robots.txt Rule Audit

Evaluates crawler-specific `Disallow:` and `Allow:` directives with wildcard matching.

Meta Robots & X-Robots-Tag

Scans HTML header tags and HTTP response headers for noindex and nofollow directives.

Auditing Webpage Crawlability & Indexability…

Fetching host robots.txt rules, tracing HTTP redirect chains, and analyzing meta directives.

Crawlability Access
Indexability Status
HTTP Status Code

200

Bot Tested
Googlebot
Crawlability & Indexability Decision Tree Evidence
Crawlability Control (Bot Payload Access)

Determines whether search crawlers can fetch the webpage file from the server.

Allowed by robots.txt and server returning HTTP 200.
Indexability Control (Search Results Eligibility)

Determines whether search engines are allowed to index this page in search results.

No noindex directives detected in HTML meta or HTTP headers.
Host robots.txt Directive Evaluation
robots.txt File Fetch: Found (200 OK)
Matched Directive for Googlebot: None
Rule Evaluation Result: Path allowed
Discovered Sitemaps in robots.txt
  • No Sitemap directives discovered in robots.txt
HTML Robots Meta & HTTP X-Robots-Tag Audit
Raw Directives Detected:index, follow
Has `noindex` Block: No
Has `nofollow` Directive: No
Redirect Chain & Canonical Link Tag
HTTP Redirect Chain
  • No redirects. Target loaded directly.
Canonical Link Tag Evaluation
Has ``: Yes
Canonical Target URL: None

What is Website Crawlability?

Crawlability describes a search engine crawler's ability to access and scan content pages on a website. If a page is blocked via robots.txt, 500 server errors, or noindex tags, search bots will fail to crawl and index it.

Frequently Asked Questions (FAQ)

Crawl budget is the number of pages Googlebot crawls on your site during a specific timeframe.

Remove unneeded Disallow directives from robots.txt, fix 500 server errors, and eliminate broken redirect chains.