We pointed the site crawl at 45squared.com, the agency site of the company that builds SEO Score API. It took 20.6 seconds to read 44 pages. It found one broken link, one redirect that five blog posts still point at, three pairs of pages with the same title, and 17 pages in the sitemap that no other page links to.

None of that shows up in a single-page audit, which looks at one URL. These problems are between the pages.

This is our own site, so every number below is real and we can publish it without asking anyone.

The crawl

curl -X POST "https://seoscoreapi.com/crawl" \
  -H "X-API-Key: YOUR_KEY" -H "Content-Type: application/json" \
  -d '{"url": "https://45squared.com", "max_pages": 100}'

The summary that came back:

{
  "pages_crawled": 44,
  "page_cap": 100,
  "truncated": false,
  "ok_pages": 42,
  "broken": 1,
  "redirects": 1,
  "duplicate_titles": 3,
  "duplicate_h1": 3,
  "orphan_pages": 17,
  "median_response_ms": 1171,
  "sitemap_urls": 42,
  "blocked_by_robots": 0,
  "rate_limited": 0,
  "seconds": 20.6
}

truncated: false matters. The crawl stopped because it ran out of site, not because it hit the cap, so these are the whole site's numbers.

Finding 1: a 404 that one post still links to

{
  "url": "https://45squared.com/common-wordpress-mistakes-and-how-to-fix-them/",
  "status": 404,
  "linked_from": [
    "https://45squared.com/the-most-common-mistakes-made-as-a-new-small-business-owner/"
  ],
  "in_sitemap": false
}

An old post was removed. Another post still links to it. The sitemap is clean, so a sitemap check would have passed.

The useful field is linked_from. A list of 404s tells you what is gone. The page that links to it tells you where to make the edit. Here that is one link in one post: point it at a live article or remove it.

Google's guidance is that a 404 for a page that is really gone is fine and does not hurt the rest of the site (Google Search Central: HTTP status codes). The cost is to the reader who clicks it, and to the post that sends them there.

Finding 2: one redirect, linked from five posts

/build returns a 301 to the homepage. Five blog posts link to /build.

It works, so nobody notices. But each of those links was written to send a reader to a specific page about building a site, and each now lands on the homepage instead. Either the page should come back, or the five links should point at the page that replaced it. A redirect is the right fix for outside links you cannot edit. For your own links, fix the link.

Finding 3: 17 orphan pages

An orphan here means a URL that is in the sitemap but that no crawled page links to. Seventeen is a lot for a 44-page site, and they fall into three groups:

Group Count What to do
Tag and category archives 13 Link them from posts, or take them out of the sitemap
Location landing pages (/chicago-web-design/, /west-michigan-web-design/) 2 Link them from the navigation or the footer
Newsletter plugin pages (?mailpoet_page=...) 2 Remove from the sitemap and set noindex

The two location pages are the ones that cost money. They exist to rank for a city, and nothing on the site points at them. Google finds pages mainly by following links, and says so in its own documentation on how links help it discover pages. A sitemap entry gets a page crawled. Internal links are what tell Google the page matters.

The plugin pages are a different problem: utility pages a newsletter plugin added to the sitemap by itself. They are not broken. They just should not be there.

Finding 4: tag and category pages with the same title

Three pairs of pages share a title and a main heading. Two of the pairs are a tag and a category with the same name:

  • /category/marketing/ and /tag/marketing/
  • /category/security/ and /tag/security/

Each pair lists nearly the same posts under the same title. The fix is to keep one taxonomy for a topic. If a category already exists, the tag with the same name adds a second URL for the same list and nothing else.

Finding 5: four slow pages

Four pages took more than 1.5 seconds to respond, with the contact page slowest at 2.0 seconds. The median for the site was 1.17 seconds.

This is server response time for the HTML, measured from our crawler. It is not a Core Web Vitals score. Google's own guidance treats a server response under 0.8 seconds as good (web.dev: Time to First Byte), so a median of 1.17 says the whole site has room, not just four pages. To see what a slow page costs a real visitor, run a standard audit on it; that includes Core Web Vitals.

The fix list, and what the crawl did not tell us

Five findings turn into a short list, in the order we would do them:

  1. Edit one post to fix or remove the link to the 404.
  2. Point the five /build links at the page that replaced it.
  3. Link the two location pages from the footer.
  4. Take the plugin pages out of the sitemap.
  5. Merge the duplicate tags into their categories.

The first two are possible in minutes only because the crawl names the pages to edit.

What the crawl cannot do:

  • It does not render JavaScript. It reads the HTML the server returns. A link that only exists after a script runs is not followed, so a site built that way will show more orphans than it really has.
  • It does not score pages. It records status, timing, title, heading, canonical and links. It will not tell you a title is weak, only that two pages share one.
  • An orphan is relative to this crawl. A page linked only from a page we did not reach counts as an orphan. On this site the crawl was complete, so the 17 are real.

Running it on your site

A crawl is included on Pro (25 pages) and Ultra (100 pages), and costs one Deep Audit credit on any other plan. Every crawl also gets a site map page you can send to a client.

The order that works: crawl first to find which pages have problems, then run a Deep Audit on the pages that earn money.

Get an API key or read the site crawl guide.