Tech for Retail 2025 Workshop: From SEO to GEO – Gaining Visibility in the Era of Generative Engines

Back to blog

SEO crawl budget: when volume becomes the problem, and what to decide per URL family

SEO

Discover Incremys

The 360° Next Gen SEO Platform

Request a demo
Last updated on

26/9/2026

Chapter 01

Example H2
Example H3
Example H4
Example H5
Example H6

On a large catalogue, the point is not to be crawled but to be crawled usefully: to bring the robots to the right URLs, at the right pace, without trapping them in low-value areas. That still leaves establishing that your site is in that situation.

 

Does crawl budget apply to my site?

 

The question is settled before any project: a crawl optimization launched on a site that does not need one produces nothing. Two working orders of magnitude serve as a first grid: beyond roughly one million unique pages whose content changes moderately — in the order of once a week — and, from roughly 10,000 unique pages, when the content changes very fast, daily. These are markers, not admission thresholds. A third criterion is read in Search Console: a large volume of URLs listed as “discovered, currently not indexed”, that is, known and not yet crawled. “Crawled, currently not indexed” does not fall under the same diagnosis.

The exposed profiles are always the same: e-commerce sites (facets, sorting, pagination, variants), marketplaces (highly dynamic inventory), media sites (freshness) and directories (volume and duplication). On that type of site, a structural drift can create tens of thousands of “parasitic” URLs that capture crawling at the expense of the categories and products that matter.

Below those volumes, the honest answer is: often less, but not never. A slow or unstable site, or one that generates a lot of parameter URLs, can suffer from inefficient crawling at a few thousand URLs. But if your pages do not appear and the site is small, the problem is rarely volume: it is a basic technical check — an over-broad directive, an inconsistent canonical, an unexpected HTTP status — and it is the technical SEO audit that brings it to light, with its checks on directives, statuses, canonicals and rendering.

 

What Google can crawl, and what it chooses to process

 

Crawl budget refers to the set of URLs a search engine can and wants to crawl on a site. Four stages can be distinguished within it, and confusing them leads to misreading the reports:

  • Discovery: Google finds the URL (links, sitemap, history).
  • Crawling: Googlebot fetches the resource.
  • Rendering: JavaScript processing and construction of the interpreted page, where needed.
  • Indexing: analysis, consolidation, then potential inclusion in the index.

A crawled URL is therefore not an indexed URL, and that nuance prevents false emergencies. On JS-heavy sites, the “rendering” cost becomes a multiplier of the crawling cost: even when the URL is reached, it can be “expensive” to process, which mechanically reduces the coverage that is possible. Another structuring point: the budget is defined per hostname; www.example.com and code.example.com do not have the same budget. For multi-subdomain architectures, that is an SEO design decision in its own right. How a URL enters the queue and which paths the robots follow belong to SEO crawling.

 

Capacity and demand: why adding servers is not enough

 

The budget results from two components that do not offset each other. Capacity depends on the number of parallel connections the robot can open and on the delay it imposes between two fetches, so as not to overload your servers: that is where the “crawl delay” vocabulary attaches. Demand reflects the engine's interest in revisiting your URLs — size, freshness, perceived quality, relevance, authority. The consequence prevents the most expensive decision: if demand is low, Google crawls less, even when the server could take more. Adding infrastructure to solve a problem of perception produces nothing. Different robots also have distinct demands, which explains frequency gaps between page types.

 

The signals that make demand vary

 

Three engines drive it: perceived inventory, authority and freshness. On large sites, perceived inventory is often the most actionable lever: too many duplicated, deleted or “unwanted” URLs waste time and can lead the systems to judge that the rest of the site is not worth going through. Conversely, certain site-wide events — a migration, typically — trigger a temporary surge in demand, the time it takes to reprocess everything. That is a window, not an acquired right: if the migration creates redirect chains and duplicates, you increase the processing cost at the worst possible moment.

 

Recognizing a genuine crawling problem

 

Before blocking anything, you have to establish that there really is a crawling problem, and not one of perceived quality, architecture or intent. The three look alike in the reports and do not call for the same decisions. A diagnosis made too quickly produces blocks that change nothing.

 

Reading the indexing states without getting them wrong

 

The most useful reading goes through the indexing reports — valid, excluded, errors — and through the two states “discovered, currently not indexed” and “crawled, currently not indexed”, which do not describe the same thing. The first designates URLs that are known and not yet crawled. The second designates pages that Google has crawled and chosen not to index: it is not proof of an insufficient crawl budget, and it is cross-checked with recrawl delays, logs and crawl statistics before any conclusion. Two false positives recur: confusing “not indexed” with “technical problem”, and treating every discovered URL as a URL to be saved. On a large site, the right question is not how to get a URL indexed, but: is this a URL that should exist in the indexing strategy? Answering no to it removes more projects than it creates.

 

The two symptoms that count

 

First symptom: pages discovered but not crawled, or crawled too late. Two causes dominate — the page is too deep or badly linked, or it belongs to a set whose perceived value is low, typically facet combinations. Excessive depth acts as a priority penalty: the further a URL sits from the home page and the hubs, the more expensive it becomes to reach and the fewer consistent internal signals it receives. Second symptom: content that changes often — price, stock, availability — without being revisited at a consistent pace. That is a freshness deficit: the engine re-crawls to capture changes, but only if it perceives that this is “worth the cost”. The signals that reduce that perception are structural: massive duplication, slow templates, server instability, or “noise” (endless URLs). The stake is direct when visibility concentrates on few positions: the click-through rate on the first organic position (desktop) is 34% (SEO.com, 2026), a benchmark collected in the SEO statistics. Freshness that is badly reflected degrades clicks and trust before it costs rankings.

 

Segment before deciding

 

On thousands, sometimes millions of URLs, you do not steer at page level but at the level of the template and the URL family: that is what makes a decision to block or to open defensible. Segment at minimum the categories and subcategories, the product pages, the facets, filters, sorting and parameters — the main source of drift — the editorial content and the technical pages.

For each segment, then compare five dimensions: (1) volume of URLs generated, (2) expected indexability, (3) crawl signals, (4) performance — TTFB, stability — and (5) business value. On a catalogue, “important for SEO” must stay aligned with “important for the business”: cross organic performance, conversion, margin, availability and seasonality, so as not to over-invest in URLs with a low contribution.

The diagnosis then isolates the URL families that meet at least one of these four criteria:

  • they are crawled frequently, but are not indexed or are excluded on a recurring basis;
  • they have no organic traffic and answer no searched intent;
  • they introduce duplicate content (close variants, parameters, pagination);
  • they degrade crawl capacity through slowness, errors or redirects.

The grid below sums up what is expected per family and the default decision.

URL family Expected indexability Signal that should raise a flag Default decision
Categories and subcategories Indexable and priority Rarely crawled despite demand Reduce depth, strengthen internal linking
Product pages Indexable, long tail Discovered but not crawled, in bulk Handle in batches from a hub
Facets, filters and sorting A selection only The number of URLs grows, the traffic does not Decide: SEO-first or UX-only
Editorial content Indexable Updates not reflected Check the perceived freshness
Technical pages Not indexable Crawling of areas with no search value Prevent discovery, not indexing alone

 

Where the budget is lost

 

Waste concentrates in three families of causes, and each one calls for a different remedy. Taking them in this order avoids blocking first and understanding afterwards.

 

Facets, parameters and sorting: the SEO-first or UX-only decision

 

Faceted navigation can produce an endless number of URLs — filter combinations, sorting, pagination. It is the classic robot trap: the engine keeps discovering more URLs, much of whose content is redundant, hence wasted crawling and an unfavourable trade-off on the pages that really matter. The right approach is not to “block everything” but to decide which combinations deserve indexing — strong search demand, a structuring range, seasonality — and to neutralize the rest through consistent rules: directives, canonicals, URL patterns, internal linking. That decision is made once, in writing: which facets are “SEO-first” — indexable, present in the sitemap, reinforced by internal linking — and which must stay “UX-only”. Without an explicit ruling, the rule rebuilds itself at the first release, always in the direction of opening up.

 

Worthless pages, duplication and the noindex trap

 

On a catalogue, internal search, the basket, the login or empty pages become crawl vacuum cleaners as soon as they are reachable by crawlers: the remedy is to close the areas that are not important to the engine, even when they are useful to the visitor — sorting variants, infinite scroll that duplicates information already reachable through links. One trap: avoid using noindex as a “crawl-saving solution”. If Google has to crawl the URL to see the noindex, crawl time is consumed anyway; it is the block in the crawl file that prevents the fetch. On the duplication side, consolidate or eliminate, so as to concentrate crawling on unique content rather than on unique URLs — and be wary of canonicalizing “everywhere”, which makes lost pages get reprocessed when the internal linking pushes towards URLs that canonicalize elsewhere. Three levers: reduce near-identical categories, limit variant pages with no search demand, rewrite templated content that is too close.

 

Redirect chains and instability: the hidden cost

 

Long redirect chains have a direct negative effect on crawling. They appear after a migration, a change of URL rules, products redirected in cascade or a late standardization. Each hop consumes a request and processing time; multiplied by thousands of URLs, that degrades the coverage of the genuinely useful pages that return 200. Capacity, for its part, follows “crawl health”: it rises when the site answers fast and steadily, it falls as soon as it slows down or returns server errors. The slower the server and the rendering, the smaller the “useful window” left to cover a large inventory. For pages deleted for good, return 404 or 410 rather than blocking: a 404 is a strong signal not to re-crawl, whereas a blocked URL stays in the queue longer. And soft 404s must be eliminated, because they carry on consuming crawling.

 

Crawl delay: what it really is

 

The expression covers two things. The Crawl-delay directive, interpreted by some robots through the site's crawl file, has to be treated as a non-universal mechanism: other robots respect it, but Googlebot ignores it. Setting it therefore does not slow Google down, as it regulates its pace on the server's response — observed capacity, errors, latency. The useful notion behind the expression is the delay between two requests, a component of the capacity limit the robot modulates so as not to overload the server. To write or fix those access rules, the syntax and use cases of the robots.txt file are handled separately.

Two scenarios look alike — less crawling — and do not have the same remedies. Deliberate limitation: you have blocked areas or reduced the accessibility of certain URL families, which can be desirable if those URLs are useless. Throttling: Google slows down because it detects slowness or server errors (5xx, instability), and therefore lowers capacity. One signal settles the matter: if the URL inspection tool returns a message of the “Hostload exceeded” kind, the limit is on the infrastructure side. That message is not the only indicator: response times that degrade under load, 5xx errors correlated with crawl peaks or measured saturation document the same limit. It is on those observations, and not on a single message, that an infrastructure reinforcement is decided.

Rather than throttling blindly, four alternatives hold: reduce parasitic URLs, stabilize the responses, improve rendering efficiency, keep the sitemaps up to date with a reliable last-modified date. You protect the infrastructure by making crawling more efficient, not by degrading the overall capacity for discovery. A sequencing reminder goes with it: de-index first, block second. A page blocked before it has been de-indexed can no longer leave the index, for want of a recrawl to read the directive.

 

The action plan at volume

 

Three workstreams run in parallel, and none replaces the others: steer crawling towards what counts, reduce what does not have to be crawled, and bring down the cost of each fetch.

 

Concentrating crawling on the strategic pages

 

At high volume, internal linking works as a prioritization system. High-impact pages — core categories, high-margin products, best sellers, pillar content — must receive more links: navigation and hubs, contextual links, “top categories” blocks, links from strong pages. Conversely, limiting links to non-indexable URLs avoids creating crawl dead ends. Reducing depth does not mean pulling everything up into the header: the patterns that hold at scale are hub pages per universe, lateral links, accessible pagination and crawlable HTML link blocks. A business page that sits too deep becomes a natural candidate for crawl delay, especially if the site generates a lot of facet URLs.

 

Reducing the useless crawled surface

 

Three moves. Prevent the discovery of what does not have to be crawled: UX-only areas are not injected into the crawlable internal linking, and a block is justified when crawling brings no value. Standardize the parameters — order, case, separators, tracking — so that the same page does not present itself under ten addresses. And treat the cause rather than the symptom, because one reservation holds the reasoning together: blocking URLs can reduce their processing by other systems, and the “freed budget” is not automatically reallocated, unless you were already at the serving ceiling. A block is therefore never announced as a crawl gain on the strategic pages: it is the explosion of URLs itself that has to be stopped.

 

Cleaning up the redirects and stabilizing the response

 

The rule is simple to test against: one redirect at most between an old URL and its final destination, ideally zero for internal links. Standardizing the versions — http and https, www and non-www, trailing slash — removes variants that are useless to crawl. Temporary redirects that linger on SEO destination pages keep uncertainty alive and cause pointless re-crawls: if the destination is permanent, switch to a permanent redirect and update the internal linking to point straight at the final URL. On the server side, a “performance” backlog is a crawl backlog: fewer errors and less latency means more capacity. On JavaScript architectures, reducing the rendering cost — useless bundles, excessive hydration — keeps it from becoming the bottleneck. One last safeguard: if your pages depend on rendering, avoid blocking CSS, JavaScript and critical images. Block the areas you do not want crawled, not the resources needed to analyse the pages that count.

 

Steering over time and avoiding regressions

 

On large sites, the difficulty is not finding anomalies, but deciding what to fix first and how to measure the effect. A minimal dashboard is enough, provided it correlates four families of signals: crawling (trends, server errors, redirects), indexing states (valid, excluded, discovered or crawled without indexing), technical performance (TTFB, stability, slow templates) and value (impressions, clicks, conversions), so as to measure the effect on the business areas.

Regressions rarely come from bad intentions, but from a release that generates new parameter URLs, multiplies near-identical pages through template duplication, introduces redirect chains, changes depth, or degrades latency and stability (5xx spikes). After every major release, check those five points on the most crawled and most business-critical templates. A single badly framed facet rule can recreate a useless “perceived inventory” within days.

A full analysis is relaunched at four moments: a redesign (URLs, templates, JS, performance); a migration or a change of rules (redirects, canonicals, sitemap); a seasonal peak; a massive addition of products or categories. Between those moments, tracking, release after release, the volume of URLs generated per family — so as to catch a badly framed facet rule before it fills the inventory — is precisely the task covered by the audit and mapping module.

 

FAQ on crawl budget and crawl delay in SEO

 

What is the definition of crawl budget in SEO?

 

It is the set of URLs an engine such as Google can and wants to crawl on a site. It combines a capacity, limited by what your server can take, and a demand, which reflects the interest in revisiting your pages. Crawling does not guarantee indexing: after the crawl, pages are evaluated and consolidated before any inclusion in the index.

 

What is crawl budget for and when does it become limiting?

 

It becomes limiting when high-value URLs — categories, products, pillar content — are not crawled fast enough, because the robot spends its time on parasitic URLs (facets, parameters, duplication, redirects) or because capacity is reduced by slowness and 5xx errors. The orders of magnitude: roughly one million unique pages updated weekly, or roughly 10,000 pages updated daily.

 

How do you know whether crawl budget is a problem on your site?

 

Five signals recur: a high volume of URLs listed as “discovered, currently not indexed” in Search Console, a delay before new pages appear, prices, stock or content updated without being reflected, crawling concentrated on parameters, sorting and technical pages, and a rise in 5xx errors or latency. None is enough on its own: it is their accumulation on the same templates that signs the problem.

 

How do you optimize without losing important pages?

 

By proceeding through segmentation and rules: first define which URL families must be indexable, then align internal linking, sitemap and directives on that decision. Avoid irreversible global actions — a broad block — without a precise inventory beforehand. For deleted pages, prefer a 404 or a 410 to a block, which leaves the URL in the queue for longer.

 

What is crawl delay in SEO?

 

It is the idea of slowing down how often a robot passes between two requests. With Google, the pace depends above all on crawl capacity, defined in particular by the parallel connections and the delay between two fetches, adjusted so as not to overload the server. That delay is therefore not a setting you impose, but a consequence of what your infrastructure returns.

 

Does the crawl-delay directive in robots.txt really improve SEO?

 

Not as a main lever. Googlebot purely and simply ignores the directive — other robots respect it, it does not — and adjusts its pace according to server health: latency, errors, stability. The reliable lever remains reducing useless URLs and improving performance, so as to raise both capacity and crawl efficiency.

 

Can duplicate pages reduce the crawl frequency of valuable pages?

 

Yes. When many known URLs are duplicated or unwanted, perceived inventory degrades and crawl time is wasted on pages that bring nothing. The answer is to consolidate or eliminate the duplication, so as to concentrate crawling on unique content rather than on unique URLs.

 

Why do redirect chains consume so much budget?

 

Because a chain imposes several requests and several rounds of processing to reach a single destination. Long chains have a negative effect on crawling, and on a large catalogue the cumulative impact quickly becomes major. The acceptance rule is one hop at most, and zero for internal links, which must point to the final URL.

 

My site has fewer than 10,000 URLs: should I worry about it?

 

Often less, but not never. If the site is slow, unstable, or generates many useless URLs through parameters, you can observe inefficient crawling at that volume too. That said, the most significant gains appear on large or very dynamic sites, according to the orders of magnitude recalled above.

 

Should filters and facets be blocked, or left crawlable on a large e-commerce site?

 

Neither block everything nor open everything. Choose which combinations have search value and a clear intent — those are SEO-first and indexable — then prevent the combinatorial explosion of the rest, which stays UX-only. A block in the crawl file suits URLs that are not important, such as sorting variants; noindex, for its part, does not prevent crawling.

 

How do you prioritize fixes on a catalogue of hundreds of thousands of pages?

 

By combining segmentation by template, business value (traffic, conversion, margin, seasonality) and crawl and indexing signals. Look first for the red zones: templates crawled often but slow, URL families that generate a lot of duplication, and the redirects or errors that recur from one release to the next.

 

Continue reading

 

  • You need to quantify what your URL families actually consume before committing to a blocking project: the log analysis gives the visit frequency per section, the statuses actually returned and the share of wasted requests.
  • The catalogue itself has become the subject, and you have to arbitrate ranges and variants rather than crawling alone: the e-commerce SEO audit handles categories, product pages, facets, stock-outs and expired products.

Discover other items

See all

Next-Gen GEO/SEO starts here

Complete the form so we can contact you.

The new generation of SEO
is on!

Thank you for your request, we will get back to you as soon as possible.

Oops! Something went wrong while submitting the form.