How to handle pagination and faceted navigation
Every shop, directory or listing site I have worked on hits the same wall. You launch with a few hundred products, someone adds filters for colour, size, brand and price, and a year later Google is crawling 400,000 URLs for a site with 3,000 real pages. Rankings for the categories that matter get diluted, new products take weeks to be discovered, and the crawl stats report in Search Console looks like a landfill.
This tutorial is for site owners and operators who run a catalogue, a directory, a blog archive or any list that splits across pages and can be filtered. I write from the operator side, running several sites from Singapore. By the end you will know which paginated and filtered URLs deserve to be indexed, which should be kept out, and how to prove it worked using your own logs.
The correct setup depends on whether your filtered pages match things people actually search for. “Red running shoes” probably does. “Red running shoes sorted by price descending, page 4” does not.
What you need
- Search Console access for the property (free), plus the Crawl stats report under Settings
- a crawler: Screaming Frog SEO Spider (the free version caps at 500 URLs, a paid licence lifts that) or Sitebulb
- access to your server or CDN access logs, or at least a way to export a week of them
- the ability to edit templates (canonical tags, meta robots, internal links) and robots.txt
- a spreadsheet for the URL pattern inventory
- a staging copy or at least a way to test changes on one category first
- two to four hours for a mid-sized site, longer if your platform makes templates hard to change
If you are unsure how canonicals behave, read my canonical tags explained piece first. Half of this tutorial depends on it.
Step by step
Step 1: inventory every URL pattern
Action: crawl the site and group URLs by pattern, not by page. You care about patterns like ?page=2, /page/2/, ?colour=red, ?sort=price_asc, ?sessionid=, and combinations of them.
Export the crawl, then count patterns from the command line:
# urls.txt is a one-column export of discovered URLs
grep -o '?[^ ]*' urls.txt | tr '&' '\n' | cut -d= -f1 | sort | uniq -c | sort -rn | head -30
Expected output: a ranked list of parameter names. On most shops the top five are sort, page, colour, size and price.
If it breaks: if your site uses path-based facets (/shoes/red/size-9/) there are no parameters to count. Split on slashes instead and count the segments at each depth.
Step 2: classify each pattern as index, crawl or block
Action: for every pattern, decide one of three outcomes. I use this table in the spreadsheet:
- index: has search demand, unique content, and a stable result set (single-facet pages like brand or colour on a big category)
- crawl but do not index: useful for users and for link discovery, but no search demand (sorting, most multi-facet combinations)
- block: infinite or pointless spaces (session IDs, calendar widgets, price sliders with arbitrary values)
Check demand before deciding. Pull queries from Search Console’s Performance report filtered by page, and test the facet phrase in your keyword tool. If nobody searches for it, it does not get indexed.
Expected output: a sheet with one row per pattern and a decision.
If it breaks: if demand is unclear, default to crawl but do not index, and revisit after 60 days. It is easier to open the gate later than to clean out a bloated index.
Step 3: make paginated pages self-canonical
Action: each page in a series (/category/?page=2) should canonicalise to itself, not to page 1. Pointing every page at page 1 tells Google to ignore the products only listed on page 2 and beyond, and those products then depend on other internal links to be found.
Google’s own guidance on this is in its pagination and incremental page loading documentation. It also confirms that Google no longer uses rel="prev" and rel="next" as an indexing signal, so do not spend template time on them.
Your template should output something like:
<link rel="canonical" href="https://example.com/shoes/?page=2">
Expected output: view-source on page 2 shows a canonical equal to its own URL, with no tracking parameters.
If it breaks: check for plugins that inject a second canonical. Two conflicting tags usually mean Google picks for you. Crawl and filter for pages with more than one canonical element.
Step 4: use real links for pagination
Action: pagination must be plain <a href> links. “Load more” buttons and infinite scroll that only fire JavaScript events leave Googlebot with a single page of products. If you want infinite scroll for users, keep crawlable paginated URLs underneath it.
Expected output: with JavaScript disabled in Chrome DevTools, you can still click through to page 2, 3 and 4.
If it breaks: in Search Console, run URL Inspection on page 3 and view the rendered HTML. If the product links are missing, the content loads after interaction and Google is not seeing it.
Step 5: canonicalise or noindex the filter combinations
Action: for the “crawl but do not index” group, you have two honest options.
- canonical to the unfiltered parent: good for sort orders and view toggles where the content is the same set reshuffled
noindex, followvia meta robots: good for multi-facet combinations where the result set is genuinely different but not worth ranking
Do not mix the two on the same URL. A canonical to the parent plus noindex sends contradictory signals. Pick one per pattern and write it in the sheet.
For the “index” group, give those pages a unique title, an H1 that matches the facet, and a short intro paragraph. A filter page with no unique copy is the same thin page I cover in how to audit and fix thin content at scale, just generated by software.
Expected output: sort and multi-facet URLs return a canonical to the parent or a noindex tag. Index-worthy facets keep a self-canonical.
If it breaks: if a noindexed page is still in the index weeks later, check that it is not also blocked in robots.txt. Google cannot see a noindex on a page it is not allowed to fetch.
Step 6: block the endless spaces in robots.txt
Action: for patterns that create near-infinite URLs, use robots.txt. Google’s guidance on managing faceted navigation lists robots.txt disallow rules, canonicals and URL fragments as the main tools, and it is clear that robots.txt is the one that saves crawl budget because blocked URLs are not fetched at all.
User-agent: *
Disallow: /*?*sort=
Disallow: /*?*sessionid=
Disallow: /*?*price_min=
Disallow: /*?*price_max=
Wildcards work in Googlebot’s robots.txt parsing and are part of the standard in RFC 9309. Test every rule before you publish, using Search Console’s robots.txt report.
Expected output: URLs matching those patterns show as blocked by robots.txt in a crawl, and your index-worthy facets do not.
If it breaks: a rule that is too broad is the classic failure. Disallowing /*?* blocks pagination as well. Run your crawler with the new robots.txt applied and confirm that page 2 of your biggest category is still reachable.
Step 7: fix internal links so you stop creating the mess
Action: robots.txt and canonicals clean up after the problem, but internal links cause it. Remove template links to sort options and arbitrary facet combos. Consider submitting filters through a form POST or as fragments (#colour=red) when you do not want them as crawlable URLs. The rules from step 2 should also decide what appears in your XML sitemap: index-group pages only, never filtered duplicates.
Expected output: your next crawl finds far fewer parameter URLs, because the crawler no longer discovers them from category pages.
If it breaks: a faceted nav component from a theme or plugin may render links server-side that you cannot edit. Look for a setting to render filters as buttons, or override the template in a child theme.
Step 8: handle empty and out-of-range pages
Action: a filter combination with zero products should return a 404, not a 200 with “no results found”. Same for ?page=99 on a category that has 12 pages. Soft 404s waste crawl and show up as a warning in Search Console’s page indexing report.
Quick check:
curl -s -o /dev/null -w "%{http_code}\n" "https://example.com/shoes/?page=99"
curl -s -o /dev/null -w "%{http_code}\n" "https://example.com/shoes/?colour=transparent"
Expected output: 404 for both.
If it breaks: some platforms always return 200 for search-like routes. Add a rule in the template that sends a 404 status when the product count is zero.
Step 9: verify with logs and Search Console
Action: give it two to four weeks, then compare. In your logs, count Googlebot hits by pattern:
grep "Googlebot" access.log | awk '{print $7}' | grep -c "sort="
grep "Googlebot" access.log | awk '{print $7}' | grep -c "page="
Verify the user agent claim with a reverse DNS lookup if you have any doubt, since scrapers fake it. Also watch the Crawl stats report for total requests and the share of 404s, and the Pages report for the count of “Crawled, currently not indexed” and “Duplicate without user-selected canonical”.
Expected output: sort= hits drop sharply, page= hits stay, and the indexed page count settles near the number of real pages.
If it breaks: if blocked patterns still get crawled, your robots.txt may have a typo, or the cache has not refreshed. Google generally refreshes it within a day, so a week of continued hits means the rule is wrong.
Common pitfalls
- canonicalising every paginated page to page 1: it hides deep products and quietly kills their discoverability
- blocking in robots.txt and expecting noindex to work: Google cannot read a tag on a page it cannot fetch, so already-indexed URLs stay indexed
- letting facets compete with the parent category: if “red shoes” and “shoes” both rank for the same query, you have a cannibalisation problem, and my guide on finding and fixing keyword cannibalisation covers the diagnosis
- trusting a plugin’s defaults: many SEO plugins noindex all page 2+ results by default, which can strand products
- changing everything at once: you cannot tell which change helped. Roll out per pattern and note the date
Scaling this
At 10x (a few thousand products), the manual sheet from step 2 is fine. A quarterly crawl is enough and you can fix patterns by hand.
At 100x (tens of thousands of products), the sheet becomes rules in code. I move the index/noindex/block decision into the template layer so a new facet cannot ship without a classification. I also add a monthly log check, because developers add new parameters without telling anyone. If you run several storefronts or regional versions, testing how each renders from different locations and sessions matters, and the multiaccountops blog covers the operational side of managing many accounts and environments cleanly.
At 1000x (marketplace or aggregator scale), crawl budget is a real constraint rather than a theory. You need log analysis on a schedule, a curated set of indexable facet pages generated from search demand data, and a proper allowlist approach: block by default, open only what has been validated.
One related warning: when a site restructures URLs at any scale, filter and pagination patterns are the first thing to break. If a migration is coming, read what happens to your links when you change domain before touching redirects.
Where to go next
- canonical tags explained: the full rules for when a canonical is honoured and when Google overrides it
- how to find and fix keyword cannibalisation: for when your facet pages and category pages fight over the same queries
- how to audit and fix thin content at scale: for deciding what to do with the facet pages you keep indexed
You can find everything else in the blog index.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-02.