← all articles

Index bloat explained: why too many indexed URLs hurt

Index bloat is what happens when a search engine holds far more of your URLs than you ever meant it to. Your site might have 400 pages worth ranking, but Google has 12,000 URLs on file, and most of them are filter combinations, tracking-parameter copies, internal search results and old test pages. The good pages are buried in the pile.

I care about this because it is one of the most common quiet problems I see on sites that have been running for a few years. Nothing breaks. No error email arrives. Rankings just drift, new pages take longer to show up, and nobody can say why. If you are new to this, this article gives you the mental model, the mechanism, and a short list of things to check first.

What it is

Index bloat means a site has many more indexed URLs than it has pages that deserve to be indexed. The “index” here is the database a search engine builds from the pages it has crawled and decided to keep. When a URL is in that database, it can appear in search results.

The key word is “deserve”. A page deserves a place in the index when it offers something a searcher might want and that is not already offered by another URL on your site. Bloat is the gap between what you want indexed and what actually is.

A few things worth separating:

  • indexed: Google has crawled the URL, processed it and stored it, so it is eligible to show in results.
  • crawled: Google has fetched the URL, but it may still decide not to store it.
  • discovered: Google knows the URL exists, often from a link or sitemap, but has not fetched it yet.

Bloat is specifically about the first one. A big pile of crawled-but-not-indexed URLs is a related problem, but it is not bloat in the strict sense. It is more of a warning sign that you are producing pages Google does not find useful.

How it works

Search engines find URLs by following links, reading sitemaps and revisiting URLs they already know. They do not judge whether you meant a URL to exist. If your site can generate it and something links to it, it is a candidate.

That is the whole mechanism. Your CMS or store creates URLs, often automatically, and each one is a possible entry in the index. These are the usual sources:

  • faceted navigation: a category page with filters for size, colour, price and brand can produce thousands of combinations. I wrote about this in faceted navigation and the pages that multiply, and it is the single biggest source of bloat on shops.
  • url parameters: session IDs, sort orders, tracking tags like utm_source and “?ref=” values create copies of the same page at different addresses.
  • internal search results: if your site search pages are crawlable, every query anyone types can become a URL.
  • pagination and archives: tag pages, date archives and author archives on a blog can outnumber the posts themselves.
  • thin or empty pages: placeholder pages, tag pages with one post, auto-generated location pages with no real content. I covered this in what thin content means in practice.
  • staging and test content: a dev subdomain or a /test/ folder that was never blocked.
  • duplicates across versions: http and https, www and non-www, trailing slash and no trailing slash, all serving the same page.

Google tries to clean this up itself. It groups near-identical URLs, picks one as the canonical version and usually shows only that one. Its documentation on consolidating duplicate URLs explains how canonical signals work. But that process is a best guess, not a promise. A rel=”canonical” tag is a hint, not a command, and Google can pick a different canonical than the one you declared. When the signals conflict, a lot of duplicates slip through.

There is also a crawl side to this. Every URL Google fetches costs it time, and every URL it fetches on your server costs you load. Google’s guide to managing crawl budget says plainly that crawl budget is mostly a concern for very large sites, roughly a million pages or more, or sites with 10,000 or more pages that change daily. But bloat is not only about crawl budget. It is also about quality signals and about your own ability to see what is going on.

Why it matters

Bloat does not always cause visible damage, and I will not pretend it does. Here is where it tends to hurt.

  1. it dilutes the signal. When ten URLs compete to answer the same query, links and relevance get split between them. The wrong version can end up ranking, such as a parameter URL with a messy title instead of your clean category page.

  2. it slows discovery and refresh on larger sites. If Google spends its crawling on endless filter combinations, your new articles and updated product pages can wait longer to be picked up. On big sites I have seen this show up as new pages sitting in “Discovered - currently not indexed” for weeks.

  3. it drags down the overall picture of the site. Google has said that it evaluates quality across a site, and a large share of low-value pages does not help. I treat that as a reasonable working assumption rather than a precise rule. Cleaning up thin pages has helped on sites I have worked on, but I cannot give you a number that applies to yours.

  4. it makes your data useless. When the Search Console page indexing report shows 40,000 indexed pages and 300 of them matter, you cannot tell whether a drop of 2,000 is a disaster or noise. This matters most during changes, and I go into that in what happens to rankings after a site migration.

Common misconceptions

“More indexed pages means more traffic.” Not so. Traffic comes from pages that match a search and satisfy it. A page that never gets an impression adds nothing, and a thousand of them can get in the way.

“I can fix it by blocking the pages in robots.txt.” This is the most frequent mistake I see. Robots.txt controls crawling, not indexing. Google’s robots.txt introduction is clear that a blocked URL can still be indexed if other pages link to it, just without a description. Worse, if you block a page, Google cannot see a noindex tag on it, so the page stays in the index. The usual order is to let Google crawl the page, serve a noindex or a proper canonical, wait until it drops out, and only then think about blocking.

“The site: search tells me exactly how many pages I have indexed.” It does not. The count from a site: search is a rough estimate and can swing wildly. For real numbers, use the page indexing report in Google Search Console, which is explained in Google’s page indexing report help page. Compare the indexed count with the number of URLs in your sitemap. A large gap in the wrong direction is your first clue.

“Deleting pages always hurts rankings.” Removing genuinely useless pages usually does not, because they were not ranking for anything you care about. The care needed is around pages that have links or traffic, which is a different decision. I wrote about the approach in deleting pages on purpose.

How to check for it

You do not need paid tools to get started. I use these in this order:

  • Search Console, page indexing report: look at the total indexed count, then at the reasons pages are not indexed. “Duplicate without user-selected canonical” and “Crawled - currently not indexed” are the ones that point at bloat.
  • your XML sitemap: count the URLs. If your sitemap lists 500 pages and Google reports 9,000 indexed, find out where the other 8,500 live.
  • a crawler: Screaming Frog SEO Spider is free up to 500 URLs and costs about GBP 199 a year for the paid licence at the time of writing. Crawl the site and sort by URL pattern to find the parameter and filter families.
  • Search Console performance report: filter by page and look for URL patterns that get impressions but no clicks, or no impressions at all.

If you run several properties, keeping a clean separation between accounts and logins is its own discipline. The notes at multiaccountops.com are useful for that side of the work.

Once you find the pattern, the fixes are boring and effective. Use noindex for pages that should exist for users but not in search. Use canonical tags where a clean version exists. Return a 404 or 410 for pages with no purpose, or a 301 redirect where there is a true replacement. Remove junk URLs from your sitemap, since a sitemap should list only canonical, indexable pages. Stop generating the URLs at the source where you can, for example by not linking to filter combinations.

Then wait. Google processes removals over weeks, not days, and the indexed count falls slowly.

Where to go from here

If this was your first look at the topic, these are the next reads I would suggest:

Start small. Pull the indexed count, compare it to your sitemap, and look at one URL pattern. That one check will tell you whether you have a problem worth the time.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-07.

for SEOs
Tracking rankings or scraping SERPs at scale?

Rank checkers and SERP crawlers get blocked and geo-skewed fast on datacenter IPs. Singapore Mobile Proxy runs real 4G/5G mobile IPs that search engines still trust, so your position data stays clean.

see plans →
read on
More from The SEO Desk

Technical SEO, link building, content and SERP strategy, and tool reviews for people who ship growth.

browse all articles →