POST /crawl starts at one URL, follows the site's own links and its sitemap, and reports what is wrong across the whole site: broken pages, redirects, duplicate titles, pages with no heading, orphan pages and slow responses. It also produces a site map page you can share.
A standard audit and a Deep Audit both look at one page in depth. The crawl is the other half: it looks at many pages briefly, to find the problems that only show up between pages.
Starting a crawl
curl -X POST "https://seoscoreapi.com/crawl" \
-H "X-API-Key: YOUR_KEY" -H "Content-Type: application/json" \
-d '{"url": "https://example.com"}'
{
"job_id": "3f9c2a...",
"status": "queued",
"max_pages": 25,
"poll": "/crawl/3f9c2a...",
"site_map": "https://seoscoreapi.com/crawl/3f9c2a.../map"
}
The crawl runs in the background. Poll GET /crawl/{job_id} with the same key; pages_done counts up while it runs. We crawled a 25-page site in under 13 seconds in testing.
What comes back
summary gives the totals. issues groups what needs attention:
| Field | What it lists |
|---|---|
broken_links |
Pages that returned an error, each with the pages that link to it |
redirects |
URLs that redirect, where they end up, and who links to the old address |
duplicate_titles, duplicate_h1 |
Groups of pages sharing a title or a main heading |
missing_title, missing_h1 |
Pages with no title or no main heading |
orphan_pages |
Pages in the sitemap that nothing we crawled links to |
slow_pages |
Pages that took more than 1.5 seconds to respond |
noindex_pages |
Pages telling search engines not to index them |
rate_limited |
Pages the site refused with HTTP 429 even after a retry |
pages has one row per URL with its status, click depth from the start page, response time, title, heading, canonical and the number of internal links pointing at it. links is the internal link graph, in case you want to draw your own.
The linked_from list on a broken page is the part that saves time. Knowing a page is a 404 is easy. Knowing which three pages still link to it is the fix.
The site map page
Every crawl has a page at the site_map URL. It shows the totals, a diagram with the start page in the middle and each ring one click further out, the lists of things to fix, and a table of every page. Pages are colored by result, so a broken page two clicks deep is visible at a glance.
The address is unguessable and the page is marked noindex, so you can send it to a client without it turning up in search.
How it behaves on your site
- It makes plain HTTP requests, four at a time, and identifies itself as
SEOScoreAPI-Crawler. - It honours
robots.txt. A disallowed URL is listed as blocked and never fetched. - It stays on the site you gave it.
wwwand the bare domain count as one site. - If the site answers 429, it waits, retries once, and then lists the page as rate limited. It never reports a throttled page as broken.
- It stops at the page cap and tells you how many more URLs it found, so a partial crawl is never mistaken for a full one.
Limits
| Plan | Pages per crawl | Crawls per month |
|---|---|---|
| Pro | 25 | 20 |
| Ultra | 100 | 100 |
| Any other plan | 25 | One Deep Audit credit per crawl |
A crawl that fails isn't counted, and a credit spent on it is returned.
Frequently asked questions
Does the crawl run the SEO checks on every page?
No. It records status, timing, title, heading, canonical and links for each page. To score a page in depth, run an audit or a Deep Audit on it. A common pattern is to crawl first, then audit the pages that matter.
Does it render JavaScript?
No. It reads the HTML the server returns, so links that only exist after scripts run are not followed.
Can I crawl fewer pages than my cap?
Yes. Send max_pages in the request.