POST /crawl starts at one URL, follows the site's own links and its sitemap, and reports what is wrong across the whole site: broken pages, redirects, duplicate titles, pages with no heading, orphan pages and slow responses. It also produces a site map page you can share.

A standard audit and a Deep Audit both look at one page in depth. The crawl is the other half: it looks at many pages briefly, to find the problems that only show up between pages.

Starting a crawl

curl -X POST "https://seoscoreapi.com/crawl" \
  -H "X-API-Key: YOUR_KEY" -H "Content-Type: application/json" \
  -d '{"url": "https://example.com"}'
{
  "job_id": "3f9c2a...",
  "status": "queued",
  "max_pages": 25,
  "poll": "/crawl/3f9c2a...",
  "site_map": "https://seoscoreapi.com/crawl/3f9c2a.../map"
}

The crawl runs in the background. Poll GET /crawl/{job_id} with the same key; pages_done counts up while it runs. We crawled a 25-page site in under 13 seconds in testing.

What comes back

summary gives the totals. issues groups what needs attention:

Field What it lists
broken_links Pages that returned an error, each with the pages that link to it
redirects URLs that redirect, where they end up, and who links to the old address
duplicate_titles, duplicate_h1 Groups of pages sharing a title or a main heading
missing_title, missing_h1 Pages with no title or no main heading
orphan_pages Pages in the sitemap that nothing we crawled links to
slow_pages Pages that took more than 1.5 seconds to respond
noindex_pages Pages telling search engines not to index them
rate_limited Pages the site refused with HTTP 429 even after a retry

pages has one row per URL with its status, click depth from the start page, response time, title, heading, canonical and the number of internal links pointing at it. links is the internal link graph, in case you want to draw your own.

The linked_from list on a broken page is the part that saves time. Knowing a page is a 404 is easy. Knowing which three pages still link to it is the fix.

The site map page

Every crawl has a page at the site_map URL. It shows the totals, a diagram with the start page in the middle and each ring one click further out, the lists of things to fix, and a table of every page. Pages are colored by result, so a broken page two clicks deep is visible at a glance.

The address is unguessable and the page is marked noindex, so you can send it to a client without it turning up in search.

How it behaves on your site

  • It makes plain HTTP requests, four at a time, and identifies itself as SEOScoreAPI-Crawler.
  • It honours robots.txt. A disallowed URL is listed as blocked and never fetched.
  • It stays on the site you gave it. www and the bare domain count as one site.
  • If the site answers 429, it waits, retries once, and then lists the page as rate limited. It never reports a throttled page as broken.
  • It stops at the page cap and tells you how many more URLs it found, so a partial crawl is never mistaken for a full one.

Limits

Plan Pages per crawl Crawls per month
Pro 25 20
Ultra 100 100
Any other plan 25 One Deep Audit credit per crawl

A crawl that fails isn't counted, and a credit spent on it is returned.

Frequently asked questions

Does the crawl run the SEO checks on every page?

No. It records status, timing, title, heading, canonical and links for each page. To score a page in depth, run an audit or a Deep Audit on it. A common pattern is to crawl first, then audit the pages that matter.

Does it render JavaScript?

No. It reads the HTML the server returns, so links that only exist after scripts run are not followed.

Can I crawl fewer pages than my cap?

Yes. Send max_pages in the request.