A customer opens the trainers category, ticks "black", then "size 42", and finally sorts "price, low to high". To them it is one search. To the server it is a new URL with three parameters, almost identical to one that already existed. Multiply that by every category, colour, size and sort order, and a shop with a few hundred products can end up exposing tens of thousands of addresses.
Google does not know up front which of those URLs matter. If you do not tell it, it decides on its own, and it often spends its time on pages that add nothing while taking longer to revisit the ones that actually sell. Let's look at what happens and how to bring some order to it.
How one listing turns into an avalanche of URLs
Parameters are the part of the address that follows the question mark, for example ?color=black&size=42&sort=price. To a search engine, each different combination is a different URL. It gets worse because parameter order counts too: ?color=black&size=42 and ?size=42&color=black show the same thing but are two addresses.
It helps to tell three types apart. Facets (colour, brand, material) filter and genuinely change which products appear. Sorting (price, newest) only rearranges the same products. And tracking or session parameters (campaigns, identifiers) change nothing about the content. Each type deserves different treatment.
What problems they actually cause
The first is wasted crawling: Googlebot spends a limited amount of time on each site, and if much of it goes on filter variants, your important pages get revisited less often. On small shops that is rarely serious; on large catalogues it certainly is.
The second is duplicates: the same product list shows up under many URLs and Google has to pick one as canonical, which may not be the one you want. The third is cannibalisation: a filtered URL competes with your main category for the same search, and links and signals get split between the two instead of adding up.
The control tools and what each one is for
They are not interchangeable, and mixing them up is the most common mistake:
- Canonical (
rel="canonical"): states which version is preferred. It is a hint, not a command, and the URLs can still be crawled. It works well for sorting and tracking parameters, which should point to the clean URL. - noindex (the
meta robotstag or theX-Robots-Tagheader): asks for the page to stay out of search results. Google must be able to crawl the page to read it. - robots.txt: stops certain paths from being crawled and saves resources, but it is not a de-indexing command. A blocked URL that receives links can still appear in Google without content, and Google will never see a canonical or noindex placed inside it.
- Non-crawlable links: if filters are applied with buttons or forms instead of normal links, search engines do not even discover them.
That is why you should not combine robots.txt and noindex on the same URL: if you block crawling, the noindex is never read.
Which facets are worth indexing
Some combinations match real searches: "women's running shoes", "leather sofas" or "long party dresses". Those pages can rank very well because they answer exactly what people type. A facet deserves indexing if it meets several conditions: demonstrable demand (you can check it in Search Console or a keyword tool), enough products, a clean and stable URL, its own title, heading and copy, and a link from the category or menu.
The practical rule is simple: if you could defend it as a category in its own right, index it. The rest (size, price, sorting, combinations of three or more filters) are better kept out of the index and, where possible, out of the crawl too.
How to check where your shop stands today
In Search Console, the Pages report lets you look for URLs containing a question mark and see statuses such as "Duplicate, Google chose different canonical than user". Crawl stats (under Settings) show what share of requests goes to parameterised URLs. And by hand you can open a filtered URL, view the source and check which canonical and meta robots it carries.
One useful detail: Google retired the old URL Parameters tool in Search Console in 2022, so it can no longer be configured there; everything is handled on your own site.
Frequently asked questions
Is it enough to add a canonical to every filtered URL?
Not always. It helps consolidate signals, but Google may ignore it if the pages are very different from each other, and it does not stop them from being crawled. On large catalogues it is combined with non-crawlable links or robots.txt rules.
Can I block every parameter in robots.txt?
Very carefully. You could accidentally block facets you do want to rank, or parameters your site uses to load content. Test the rules against real examples before publishing them.
My shop is small. Do I need to do all this?
Probably not to that level. Even so, make sure sorting and tracking parameters point to the clean URL with a canonical and that no pointless filter gets indexed. It takes little work and saves trouble once the catalogue grows.