Tech for Retail 2025 Workshop: From SEO to GEO – Gaining Visibility in the Era of Generative Engines

Back to blog

E-commerce SEO audit: which catalogue URLs deserve to be indexed

SEO

Discover Incremys

The 360° Next Gen SEO Platform

Request a demo
Last updated on

26/9/2026

Chapter 01

Example H2
Example H3
Example H4
Example H5
Example H6

What a catalogue audit looks for, and what sets it apart

 

If you already master the fundamentals of the SEO audit, the question is no longer "why audit" but what a catalogue changes about the exercise. On a store, the slightest indexing debt — pagination, facets, inconsistent canonicals, duplication — dilutes the visibility of the pages that carry revenue. The stake lies in one decision repeated thousands of times: which catalogue URL deserves to exist, and which one deserves to be indexed. It plays out where demand concentrates — the share of global traffic going through Google is 92.96% (BrightEdge, 2024) — and it is taken template by template.

 

The catalogue produces URLs faster than it produces demand

 

A brochure site manages a few hundred stable URLs. A store lives to the rhythm of product additions, removals, variants and seasonality. The first trap is not running short of content, but letting the site produce too many URLs for a finite crawling capacity: facets, pagination, sort parameters and variants reachable through several paths create a mass of URLs that are useless to Google, with risks of duplication and wasted crawl budget.

Seasonality adds a second layer: temporary pages, seasonal products, collections renewed each season. Without hygiene rules — canonicals, redirects, controlled deindexing — the URL history inflates the index and dilutes signals. The problems are therefore no longer limited to tags to be fixed: they touch URL governance and Google's ability to prioritize the pages that generate sales.

 

Indexable does not mean priority

 

This is the dividing line of the whole exercise. A technical audit can conclude that a page is indexable; a revenue-driven audit has to conclude that it deserves to be indexed first. Indexing a combination of facets makes sense if it matches real demand and leads to products that are available and competitive. Conversely, indexing pages that show few products, unstable stock or thin margins ties up your crawl budget at the expense of strategic categories.

This reading crosses four signals that SEO alone does not give: availability, margin, conversion, average order value. The mechanisms themselves — what a canonical is, how a redirect chain reads, what a markup check validates — belong to the technical SEO audit, which sets out the rules and the acceptance criteria. What is decided here is which one you choose, for which template, and why.

 

Framing: page families, sources and crawl

 

On a catalogue, framing separates the theoretical report from a roadmap you can act on. The aim is not to audit everything URL by URL, but to cover the representative templates. A practical consequence, which also holds when the work is handed to a third party: the scope of a catalogue audit is counted in templates, not in URLs.

 

The four page families and their role

 

Start by defining the page families and their role. Category pages are transactional hubs: they structure demand and distribute authority. Product pages are conversion pages, numerous and unstable (stock, variants, reviews). Filtered pages are system pages to be governed — not to be confused with pagination, which carries the discovery paths towards deep products. Editorial content captures the long tail. The goal is not to produce more pages, but to decide which pages deserve to be indexed, which must stay crawlable without being indexed, and which must be blocked or consolidated.

Page family Role in the catalogue Default decision The signal that calls it into question
Categories and subcategories Transactional hubs, distribute authority Indexable Two subcategories target the same intent, or the assortment is too thin
Product pages Conversion pages, numerous and unstable Indexable, one reference URL per product Variants multiplied, strictly supplier descriptions, product withdrawn
Paginated pages, then filtered pages Two distinct roles: the paginated series opens access to deep products; the filter serves the buying journey Pagination: crawlable and indexable, self-referencing canonical on each page. Filters: crawlable, out of the index by default Pagination: series links unreachable, or canonical to page 1. Filters: a combination carries stable demand and a sufficient assortment
Editorial content Long tail, support for the buying decision Indexable The guide and the category compete for the same query

 

What the crawl must reveal, and what to cross it with

 

The store may run on any technology: the audit stays independent of the CMS and relies on an external crawl, from which six readings open the file:

  • the URL structure — parameters, directories, path consistency — and depth;
  • HTTP statuses and redirect chains;
  • crawl directives and sitemap quality, the sitemap having to contain only canonical URLs to be indexed;
  • canonical tags and their consistency with indexability and redirects;
  • the templates that generate duplication — pagination, facets, sorting;
  • performance signals on business pages.

This "machine" snapshot only takes on meaning once cross-referenced with Search Console and with a catalogue export: URL, category, stock status, price, margin. That is not SEO, but it is what makes prioritization possible. The last reading is read template by template: the share of users who leave a site if loading is too slow is 40 to 53% (Google, 2025), which does not weigh the same on a category as on a product page with zero impressions.

 

Architecture: depth, hubs and what internal linking pushes

 

Architecture determines how quickly engines discover the important categories and reach deep products. Aiming for key pages to be reachable in three clicks at most remains the simplest rule. On large catalogues, every unnecessary redirect and every duplicated URL consumes crawl budget: the audit looks for crawl leaks before marginal optimizations.

 

Categories as entry points, and the template that has to carry the query

 

Category pages are the best transactional entry points: stable intents, structured internal linking, signals that are easy to concentrate. The audit question is not "should every category be indexed", but: which categories carry real demand and an assortment rich enough? For subcategories, check that they do not compete with one another and that they have a listing that justifies their existence in the SERP.

The shape of the SERP then settles a trade-off you would think editorial and that is in fact structural. Transactional, and the category has to carry the query; informational, and it is the opposite, and a bare category will not hold. The audit consequence is therefore not "rewrite the page" but choosing which template you reinforce, accepting that you will not reinforce the other.

 

Native internal linking: discoverability and relevance

 

A store has powerful native linking elements: menus, breadcrumbs, "related products", "best sellers" and "recently viewed" blocks. Two checks are enough to audit them.

  • Discoverability: are these links present in the rendered HTML, and therefore reachable by robots without friction?
  • Relevance: do they reinforce the priority business pages rather than redistributing value towards sort URLs, variants or temporary pages?

On sites where internal links point massively to parameterized URLs, you see a domino effect: an explosion of crawled URLs, duplication, then the weakening of the main categories.

 

Pagination, sorting and parameters: crawlable or indexable

 

Pagination serves two contradictory goals: giving access to the depth of the catalogue, and not multiplying low-value pages. Good management rarely means blocking everything or indexing everything, but setting rules according to the size of the category, the stability of the offer and the business contribution. The risk comes from the combinations: pagination plus sorting plus facets, and one category generates hundreds of URLs. Pagination and facets are therefore not handled together — the first opens up paths, the second multiplies neighbouring listings — and the two decisions are taken separately:

  • Let Google crawl and index the paginated series: pages 2, 3, 4 carry the discovery paths towards deep products, and taking them out reduces the crawling of the catalogue.
  • Contain the indexing of what grafts itself onto the series — sorting, facets, tracking parameters — because it is the combination that creates the massive duplication, not pagination itself.

Three checks make the decision verifiable. Depth: are important products buried behind too many clicks? Internal linking: do paginated pages receive accidental internal links that amplify their visibility at the expense of the main categories? Canonicals: does each page of the series canonicalize to itself — the only configuration that holds, the one the technical audit sets out — or to page 1, which amounts to declaring that page 2 and the ones after it exist without counting? A noindex on the series produces the same impoverishment.

The same reasoning holds for parameters, with one refinement: the audit must map the parameters actually generated in internal links, not only those that are possible. Each family then receives one of three outcomes — indexable when it carries demand of its own, neutralized when it must stay reachable without competing in the index, uncrawlable when it does nothing but consume crawl budget. Sort pages almost always fall into the second: nothing requires removing them from the interface in order to remove them from the index.

 

Facets: which combination deserves a page

 

Faceted filters become a problem as soon as each combination generates an indexable URL. The audit does not remove them: it separates those that serve the user and Google, those that serve only the user, and those that produce nothing but noise. The symptom, when nothing has been settled: thousands of URLs discovered, few useful pages indexed.

 

The four criteria, and what to do when they contradict each other

 

Turning a combination of filters into a dedicated page is a selection exercise, not blind automation. Four criteria govern it.

Criterion What you look at Decision threshold If it is not met
Observable demand Impressions and queries in Search Console, internal search The combination is already searched for as such Stays a filter: it is contained, never promoted
Sufficient assortment Number of products and diversity of the listing The listing holds without emptying out Left out of the index — noindex or out of the internal linking — never canonicalized to the parent category
Stability Stock movement over several weeks The page does not go empty every fortnight Postponed: a page that empties out costs more than it returns
Business value Margin, average order value, conversion rate of the segment The segment weighs in revenue Comes after the profitable combinations, without being excluded

 

The criteria contradict one another more often than they agree, and that is where the rule is missing. Stability takes precedence over demand: a heavily searched combination on an assortment that empties out produces an out-of-stock page, bad for the buyer and for the engine alike. A profitable facet with no observable demand stays a filter. One last order applies: you decide indexability before you decide content, because enriching a page you will end up deindexing is wasted work. One clarification, because the mistake is a common one: setting a facet aside does not mean canonicalizing it to its parent category. The canonical designates identical or very similar content, whereas a thin assortment is a commercial priority, not a duplication — the parent category shows other products. A low-value facet is set to noindex, or simply left out of the internal linking.

 

Containing the rest without breaking the shopping experience

 

For the facets you do not select, the point is to preserve the journey while containing the debt. Three levers combine according to what the crawl shows: neutralizing indexing, canonicalization reserved for the URLs that genuinely serve the same listing — parameter order, case variants — steering crawling on the robots side. None of them removes the filter from the interface: the customer keeps filtering, the engine stops collecting.

That leaves the check almost everyone forgets. If your menus, breadcrumbs or "popular" blocks point to filtered URLs, you are injecting those URLs into the core of internal linking, and no indexing rule will offset that draught. The audit therefore checks the sources of internal links, not only the pages — often the most profitable fix, because it applies to a template.

 

Variants, canonicals and duplication: a single reference page

 

The same product often exists under several URLs: variants, navigation paths, parameters, campaign pages. Duplication is the costliest problem in a store because it multiplies mechanically. The audit quantifies first, then designates for each template the version that concentrates the signals.

 

Designating the reference page of a product, whether it has variants or has gone

 

The question of variants is not settled by a single rule. Some deserve a dedicated page if they match a distinct search intent; others are consolidated onto a main page to avoid cannibalization and dilution. The audit decides case by case at template level, then checks that the site applies it everywhere: URL, canonical, internal linking, sitemap. Four canonical errors show up in a crawl: a canonical pointing to a non-indexable URL (noindex, blocked, redirected); a canonical inconsistent between versions (http/https, www or not, trailing slash); canonical chains; a "default" canonical to the category when the product page has to remain the target. The acceptance criterion is a single one — alignment between canonicals, redirects, sitemap and indexability — failing which indexing becomes unstable. Product markup follows the same requirement of consistency with visible content.

A product that disappears raises the same question in reverse, in three branches. Replaced: direct redirect to the closest equivalent, and correction of the internal links. Temporarily unavailable: page kept, with alternatives. Withdrawn with no equivalent: the only case to arbitrate, and the criterion is what the page has earned — if it carries impressions, positions or inbound links, it is kept as an end-of-line page pointing to its category; otherwise it goes out as an accepted 404. Above all, remove the redirect chains.

 

Internal and external duplication: two different levers

 

Internal duplication concerns your own competing URLs — variants, facets, categories that are too close; it is handled through decisions on indexing, canonicals, consolidation and internal linking. External duplication concerns product copy reused elsewhere, starting with supplier descriptions; it is handled through uniqueness, enrichment and proof. The aim is not to write longer, but more specific to your offer. When two close categories compete for the same intent, the call belongs to the semantic audit, which gives the detection signals and the create, optimize or consolidate rule.

The remediation plan is organized into four families: rewrite the product pages that already carry impressions or revenue but remain generic; consolidate the pages that cannibalize each other; create a few dedicated pages with strong potential rather than indexing the whole filter engine; exclude from the index the low-value URL families. The priority stays business: align indexing with value, do not clean for the sake of cleaning.

 

From diagnosis to backlog: prioritize, test, verify

 

A catalogue audit is only worth what it puts into production. The difficulty is not finding anomalies — a crawl returns thousands of them — but ordering them, so that teams are not tied up on low-value tickets while the structural subjects wait their turn.

 

The order of treatment and the scoring

 

The sequence comes in four stages: secure crawling and indexing (the blockers); reduce massive duplication (parameterized URLs, facets, sorting); reinforce the structural transactional pages; industrialize the enrichment of high-potential content. Each recommendation then receives four scores: SEO impact — does it improve the indexing of business pages, does it reduce massive duplication, does it reinforce linking towards the hubs? —, effort (technical load, content, testing), risk (regressions, broken navigation) and dependencies: does it need a template rebuild, a change of URL rules, a product or marketing sign-off?

This grid avoids a classic bias: handling the easy but unprofitable fixes first while pushing back the structural subjects — facets, duplication, category linking. The order of magnitude of the gains helps make the case: the share of clicks taken by the top 3 organic results is 75% (SEO.com, 2026), while the click-through rate on page 2 of the SERPs is 0.78% (Ahrefs, 2025). The SEO statistics bring these benchmarks together with their source.

 

Testing a catalogue, and what it has to show

 

Testing is planned from the audit stage and applies to the template, not to three URLs checked by hand. Five checks are enough: the consistency between robots.txt and sitemap, the latter having to contain only canonical URLs to be indexed and to update on additions and removals; the disappearance of redirect chains and the correction of intermediate internal links; the fall in parameterized URLs injected into internal linking; the movement of the indexing reports in Search Console; the movement of speed on business pages. A 2-second slowdown pushes the bounce rate up by +103% (Hubspot, 2026): a signal to watch on the pages that carry the volume.

Three families of gains are then verified: SEO gains (impressions, clicks, positions on the targeted segments), indexing gains (share of strategic pages indexed, fall in noise) and business gains (organic conversion and attributable revenue, depending on your attribution model). The ratio between gains and cost is calculated simply — (gains – cost) / cost × 100 — provided you accept the timescales: the first effects often appear within 1 to 3 months if the recommendations are applied quickly, while some levers, popularity and link building, read rather over 6 to 12 months. It is better to enrich the 20% of product pages that generate 80% of the opportunities than to rewrite everything. Between two audits, keeping an up-to-date record of the URLs crawled and indexed by template is the task covered by the audit and mapping module: indexing debt becomes visible as it forms.

 

FAQ on the SEO audit of an e-commerce site

 

What SEO problems are specific to e-commerce?

 

They come from the scale effect: an explosion of URLs through facets, sorting and pagination, internal duplication (variants, multiple paths), external duplication (supplier descriptions), inconsistent canonicals, excessive product depth, internal linking that pushes parameterized URLs, and instability tied to stock. Performance and journey issues come on top, because a store that ranks but converts poorly loses part of the value of SEO.

 

How should pagination be managed for SEO?

 

Keep the paginated series crawlable and indexable: Google has to be able to discover deep products, and a series page taken out of the index is revisited less often. The configuration that holds is a self-referencing canonical on each page — neither a canonical to page 1, nor a noindex on the series. Then check depth, internal linking towards paginated pages, canonical consistency and the pagination plus sorting plus facets combinations: those are what you contain, not pagination.

 

What do faceted filters do and what is their SEO impact?

 

They refine a listing (price, size, colour, brand) and can generate thousands of URLs. Possible impact: wasted crawl budget, duplication, cannibalization between pages, weakening of the main categories. The answer is to select a few useful combinations to turn into dedicated pages, and to contain the rest through indexing and linking rules.

 

How do you reduce duplicate content on an online store?

 

Identify the dominant source first: supplier copy (external duplication), parameterized URLs (technical duplication), variants (internal duplication). Then prioritize: enrichment of the high-potential product pages, consolidation of the pages that cannibalize each other, parameter governance, and canonicals consistent with the sitemap and indexability.

 

How do you audit thousands of product pages without rewriting everything?

 

By segmenting. In Search Console, spot the product pages with impressions, positions close to the top 10 or a low CTR, and in analytics those that contribute to revenue. Then work by templates and by batches — same product families, same defects — rather than page by page. The aim is to improve first what carries demand and revenue.

 

Should sort pages (price, popularity, new arrivals) be indexed?

 

In most cases, no: they duplicate the same product lists with no search intent of their own. Check whether they receive significant impressions and answer stable demand. Otherwise, contain them — sorting stays available to the user — so as to avoid diluting indexing and internal linking.

 

How do you handle out-of-stock or discontinued products without losing SEO?

 

Avoid systematic 404s if the page has built up visibility and links. Three treatments depending on the case: a direct redirect to the closest equivalent when the product is permanently replaced, a page kept with alternatives when it is temporarily unavailable, an end-of-line page when it is withdrawn with no equivalent but keeps what it has earned. Also fix the internal links pointing to removed or redirected URLs.

 

What should you check on the canonicals of an e-commerce site?

 

That a single main URL is declared for each piece of content, category or product; that the canonical points to an indexable page — neither noindex, nor blocked, nor redirected; and that it stays consistent with the sitemap and the redirects. On facets and variants, check that it really serves the decision taken for the template, otherwise indexing becomes unstable.

 

Which indicators should you track after the audit to measure the impact on sales?

 

On the Search Console side: impressions, clicks, CTR, positions on the priority category and product segments, and the movement of indexing. On the analytics side: organic sessions, conversion rate, revenue attributed to the organic channel, engagement on key pages. Add technical health indicators — volume of useless URLs crawled, 4XX and 5XX errors, canonical stability — to check that the debt is falling.

 

Is an SEO audit on PrestaShop different (facets, pagination, duplication)?

 

The principles stay the same, but the risks concentrate on the multiplication of URLs by facets and parameters, pagination, mobile performance and duplication. The audit therefore has to reason by templates rather than by isolated URL, and settle the indexability of each family closely.

 

Continue reading

 

  • The diagnosis is done and you now have to build rather than fix: e-commerce SEO covers catalogue strategy, structure and commercial content.
  • The template fixes are live and you have to prove their effect on the pages that carry revenue: SEO tracking sets out the indicators, the annotations and the control of gains over time.
  • Your product pages are visible, receive traffic and do not convert: visibility is no longer the problem, and the CRO audit tackles the conversion rate.
  • You run the crawl and cross the exports yourself: the choice and the assembly of SEO audit tools.
  • You are handing the audit over and have to frame a scope counted in templates, not in URLs: what to require from a provider during an agency SEO audit, quote and acceptance testing included.

Discover other items

See all

Next-Gen GEO/SEO starts here

Complete the form so we can contact you.

The new generation of SEO
is on!

Thank you for your request, we will get back to you as soon as possible.

Oops! Something went wrong while submitting the form.