40 checks in Crawlability & Site Architecture. Automated Its subsection has a live automated checker in the engine.
| ID | Check | Passes when |
|---|---|---|
| 2.1.01 | Fetch /robots.txt (raw HTTP) | HTTP 200 |
| 2.1.02 | Ensure robots.txt is not returning 404 | Exists |
| 2.1.03 | Validate robots.txt is under 500KB | Yes |
| 2.1.04 | Validate robots.txt content-type is text/plain or text/plain; charset=utf-8 | Correct MIME |
| 2.1.05 | Check for syntax errors in robots.txt | No syntax errors |
| 2.1.06 | Identify "Disallow: /" for User-agent: * | Not allowed unless intentional |
| 2.1.07 | Ensure homepage URL is allowed by robots.txt | Allowed |
| 2.1.08 | Ensure main service pages are allowed | Allowed |
| 2.1.09 | Ensure blog pages are allowed | Allowed |
| 2.1.10 | Identify any blocked CSS files | No blocks |
| 2.1.11 | Identify any blocked JS files | No blocks |
| 2.1.12 | Identify any blocked image directories | No blocks |
| 2.1.13 | Identify any blocked /wp-content/ directories (WordPress) | Should not block core assets |
| 2.1.14 | Identify any blocked /uploads/ directory | Not blocked |
| 2.1.15 | Identify broken or disallowed JavaScript frameworks | Allowed |
| 2.1.16 | Validate formatting of Allow: and Disallow: | Correct usage |
| 2.1.17 | Check if sitemap is referenced in robots.txt | Sitemap present |
| 2.1.18 | Validate sitemap URL in robots.txt returns 200 | 200 |
| 2.1.19 | Validate sitemap URL uses HTTPS | HTTPS |
| 2.1.20 | Count number of Sitemap: entries | Reasonable (< 10) |
| 2.1.21 | Validate wildcard pattern support (*) | Valid usage |
| 2.1.22 | Validate end-of-line comments (#) | No interference |
| 2.1.23 | Validate blank lines interpretation | No malformed spacing |
| 2.1.24 | Validate no disallow for /wp-admin/admin-ajax.php (WordPress) | Not disallowed |
| 2.1.25 | Validate no conflicting Allow/Disallow pairs | No conflicts |
| 2.1.26 | Validate no-case sensitivity violations | Correct casing |
| 2.1.27 | Validate robots.txt does not exceed 10,000 lines | Yes |
| 2.1.28 | Validate robots.txt does not block JSON endpoints | Allowed |
| 2.1.29 | Identify disallow of /api/ or /rest/ if site uses JS frameworks | Avoid blocking critical APIs |
| 2.1.30 | Validate no disallow rules match canonical URLs | True |
| 2.1.31 | Identify excessive Allow: rules (>= 20) | Reasonable number |
| 2.1.32 | Identify excessive Disallow: rules (>= 20) | Reasonable number |
| 2.1.33 | Detect empty wildcard rules (Allow: *) | Avoid unnecessary wildcard |
| 2.1.34 | Validate that URL-encoded characters (%) are correct | No malformed encodings |
| 2.1.35 | Detect robots.txt cloaking (serving different robots.txt to Googlebot vs browser) | Identical |
| 2.1.36 | Validate no "Disallow: /*?" rules that break query URLs | Avoid unless intentional |
| 2.1.37 | Detect if robots.txt is served with caching headers | Cache-Control present |
| 2.1.38 | Validate robots.txt size stability (same file across repeated loads) | Stable |
| 2.1.39 | Ensure robots.txt loads consistently across user-agents | Identical for Chrome, Googlebot, Bingbot |
| 2.1.40 | Compute total robots.txt compliance composite score | >= 90% |
A Deep Audit scores every check here that applies to your page, then an AI pass of up to 150 checks. Included on Pro and Ultra, or $9 for one audit.
See plansMachine-readable: catalog totals and the full catalog as JSON (IDs, section, weight, status).