Project · Appliance e-commerce

Auditing 70,637 indexed URLs before a replatform

Marvel Refrigeration was about to move to a new platform. The index had grown to roughly 300 times the size of the actual catalogue, and nobody knew which URLs mattered. Moving in that state would have been guesswork.

Brand
Marvel Refrigeration
Sector
Appliance e-commerce
Engagement
Pre-migration technical audit
Disciplines
Technical SEO, indexation, IA
70,637URLs in the index at the start of the audit
99.7%identified as junk faceted filter combinations
0URLs moved before the redirect map was signed off

The problem

Faceted navigation on the product listing pages generated a new crawlable URL for every combination of filter. Colour, capacity, finish, installation type and price band, all of them multiplying against each other. Google had found and indexed tens of thousands of these.

On its own that is a crawl budget problem. With a replatform coming, it became a migration problem. You cannot build a redirect map for URLs you have not counted, and you cannot decide what deserves a redirect until you know what is earning anything.

The question that mattered

Not “how do we clean up the index” but “which of these 70,637 URLs has ever earned a click, a link or a ranking”. Everything else is a 410, not a redirect.

What I found

A full crawl plus Search Console export, joined on URL, split the index into four buckets.

  • Earning pages. Categories, products and a small set of filter combinations with real impression volume. These needed one to one redirects.
  • Canonical mismatches. Pages declaring a canonical that Google was ignoring, usually because the target was itself filtered or paginated.
  • Orphan facets. Indexed, crawled, no internal links pointing at them any more. Pure crawl waste.
  • Soft 404 territory. Filter combinations returning zero products but a 200 status and a full template.

FIGURE 1
Search Console coverage report showing indexed URL count before the audit
Replace with a real screenshot, 1600 x 900

What I did

  1. Built a URL inventory joining the crawl, the Search Console export and the server logs, so every URL had a crawl frequency and a performance number against it.
  2. Wrote the containment rules for the new platform: which facets stay crawlable, which get noindex, which never generate a link at all.
  3. Mapped redirects only for URLs with evidence behind them, and marked the rest as 410 so the index would shed them cleanly rather than slowly.
  4. Modelled internal link equity on the new architecture, so the pages that mattered were not suddenly four clicks from the homepage.
  5. Wrote the staging QA checklist and the launch day monitoring plan.

FIGURE 2
The URL inventory sheet: bucket, evidence, action, destination
Replace with a real screenshot or an anonymised export, 1600 x 900

The result

The migration went out with a redirect map covering every URL that had ever earned an impression, and a deliberate decision recorded for every URL that did not. No guessing on launch day, and no six month cleanup afterwards.

Add your post-migration numbers here. Organic sessions at 30, 60 and 90 days against the pre-migration baseline, indexed URL count after the index settled, and the click recovery curve. Those three charts are what turn this from a description into evidence.

What I would do differently

I would pull the server logs earlier. I had them by the time the redirect map was being built, but crawl frequency data from the start would have made the earning versus orphan split faster and more confident.

Tools used

Screaming Frog, Google Search Console, server log files, Google Sheets for the inventory join, and Python for the deduplication once the sheet passed a hundred thousand rows.

Have a migration coming up?

The cheapest time to fix a migration is before it ships. Twenty minutes on a call will tell you whether yours is at risk.

Typical reply time: within one working day