26/9/2026
Published pages that never show up, a section the robot seems to ignore, an indexing report nobody knows how to read: these symptoms all raise the same question. What does Googlebot actually do with your site, between the moment a URL exists and the moment it enters — or does not enter — the index?
What SEO crawling covers: discovery, crawling, indexing
Crawling is the non-negotiable condition of everything else: without it, Google discovers, understands and processes nothing. But it is a mechanism, not a box to tick, and understanding it stops you from treating the symptoms backwards. If your need is to run checks — verify the directives, the statuses, the canonicals, the pagination, the rendering, then sort them by impact — it is the technical SEO audit that carries them. What follows explains the mechanism those checks verify: how a URL is discovered, queued, fetched, rendered, then kept or not.
What a crawler observes of a site
To “crawl” means, literally, to scan: in an SEO context, a website crawl consists in extracting as much information as possible in order to understand the structure, check how robots reach the pages and detect anomalies that can harm visibility — fragile site structure, insufficient internal linking, duplicated metadata. That reading “from the outside” reconstructs what a robot can reach through the links and the available signals, independently of the CMS or the framework. A diagnostic crawler simulates that behaviour: it visits URLs, follows the links it meets, collects HTTP statuses, spots redirects, measures structuring elements — titles, canonicals, robots directives — and brings out the areas at risk of being poorly discovered or poorly understood.
Crawled is not indexed, and the reverse is true as well
Crawling corresponds to searching for and analysing content so that it can potentially be displayed in the results, whereas indexing consists in deciding to add (or to keep) a URL in the index. A page can therefore be visited by Googlebot without being eligible for display in the SERPs. The frequent cases are well known: a noindex directive, duplication that leads Google to keep another canonical URL, content judged of little use, technical inconsistencies.
The converse is the least understood consequence, and the most expensive: preventing a URL from being crawled does not mechanically cause it to be de-indexed — if Google can no longer reach the page, it cannot see a noindex, a 301 redirect or a 410 code.
How Google crawls: who passes, at what pace, and what it has to load
Google uses automated systems, Googlebot among them, to discover and revisit URLs. The pace depends on the perceived importance of the pages, on their freshness and on technical constraints, and it is nothing like instant: the process follows queues, with a first stage where text content is processed after discovery and crawling, then a second where rendering can lead to a re-indexing of the final content.
On the scale of the web, crawling is massive: the number of results crawled by Googlebot each day is 20 billion (MyLittleBigWeb, 2026), a benchmark collected in the SEO statistics. That does not say how many requests are devoted to your site, but it is a reminder of one reality: Google prioritizes. Your server constraints enter that trade-off — if the site regularly returns errors, 5xx in particular, or shows high latency, Googlebot can reduce crawl pressure to avoid overloading the infrastructure, which slows down how quickly updates are taken into account.
URL discovery: internal linking, external links, sitemap
Before crawling can happen, Google has to discover the URL. Three sources dominate in practice:
- internal linking (menus, contextual links, pagination): it determines the “natural” access paths and the depth of the pages;
- external links: they contribute to discovery and prioritization, and can keep a URL “alive” even when it becomes orphaned internally;
- the XML sitemap, which lists URLs to crawl as a priority or to revisit, with no guarantee of immediate crawling.
The nuance settles a great many disappointments: a sitemap does not force the robot to pass; it is a discovery and prioritization signal, not an “index now” button.
Rendering and resources: what Google has to load
Crawling a URL does not always amount to understanding a page. Robots can use JavaScript and CSS to analyse the DOM, which sometimes implies a rendering stage. If the main content or the links are only available after JavaScript has run, the analysis becomes more expensive, slower and more prone to gaps between the initial HTML and the rendered content. The safeguard is concrete: blocking resources too broadly in robots.txt (CSS, JS, fonts, critical images) can degrade the rendering, therefore the evaluation, and in the end Google's ability to interpret the page correctly.
The fundamentals that make a site easier to crawl
Three families of causes explain most failed crawls: an architecture that pushes pages away, URLs that multiply, and unstable server responses. All three can be read in the crawl.
Architecture and internal linking: reducing depth
Internal linking plays a double role: it helps to discover pages and to understand their hierarchy. The deeper a page sits — more clicks from the entry point of the site — the harder it is to reach and the more it risks being revisited rarely. A practical rule often used in audits is to aim for important pages reachable in around three clicks, relying on topical hubs and contextual links. Three points are watched first in a crawl: orphan pages, with no incoming internal link despite business interest or backlinks; non-crawlable links, whose navigation depends on complex interactions; and internal links pointing massively to non-indexable URLs (noindex, redirects), which turn the internal linking into dead ends.
URL management: parameters, facets and duplication
Parameters and facets can multiply URLs endlessly: sorting, combined filters, internal search pages, sessions, UTMs. Crawling scatters, and Google spends time on variants with no value. The key lever is not only to block, but to decide which URLs deserve to be indexable and canonical: canonicalization exists precisely to flag duplicate pages and avoid excessive crawling. With one nuance that prevents the most frequent mistake — canonicalizing a URL A to B is not a “clean” deletion method if A and B are genuinely different — and nor is a redirect, because two distinct pages do not on their own justify a 301: that presupposes a relevant replacement. With no equivalent, a 404 or a 410 is the right answer.
HTTP codes, redirects and server stability
HTTP statuses structure the robot's experience: a URL returning 200 can be processed, a 404 signals a missing resource, and 3xx codes point to another destination. The most expensive problems for crawling, especially at scale:
- redirect chains (3xx → 3xx → 3xx), which consume requests and delay access to the final content;
- temporary redirects (302) used where a 301 would be expected to stabilize a lasting change;
- 404 errors on pages that should exist — broken internal linking, faulty template, uncontrolled deletion.
For permanent deletions, a 410 code can speed up de-indexing, whereas a 301 is only justified towards a genuinely equivalent page: the presence of external links on the deleted URL is not enough to redirect it to content that is not its replacement. That leaves stability: an unstable server — 5xx spikes, timeouts — brings a drop in crawl frequency, and the more “expensive” the pages are to process, the more Google has to arbitrate.
The sitemap: what it does, what it does not do
An XML sitemap becomes genuinely useful in three typical cases: large sites, fresh content published often, deep pages that internal linking barely reaches. Conversely, on a small, perfectly linked and stable site, it mainly adds a layer of control — monitoring, checking for gaps — rather than a discovery gain. And in every case, it does not compel crawling: it flags URLs that have been added or modified, which can help with prioritization, but the final decision depends on perceived quality, on the consistency of the signals and on constraints.
An SEO-oriented sitemap is therefore not an inventory of everything that exists: it reflects the indexing strategy — URLs returning 200, indexable, canonical, genuinely useful. Three classic mistakes degrade the quality of the signal:
- including URLs that are redirected, in error, or set to noindex;
- mixing in variants (parameters, facets) when the canonical points elsewhere;
- letting “technical” URLs — internal search, baskets, accounts — contaminate the file.
Segment by type if needed (articles, categories, local pages) to make diagnosis easier, and always align sitemap, canonicals and internal linking: if your sitemap pushes a URL, but your site treats it as a variant, you are manufacturing useless crawling. The most actionable check comes next: compare the URLs submitted through the sitemap with the real state on Google's side. The “submitted” vs “indexed” gap often reveals problems more structural than isolated errors — duplication, insufficient perceived quality, canonicalization conflicts, or an architecture that produces too many variants.
Controlling robot access without degrading SEO
The crawl file tells robots which pages or which files they may or may not request. Used correctly, it acts as a safeguard to limit crawling of worthless areas — internal search, non-strategic parameters. Used too broadly, it prevents access to business directories or to resources needed for rendering, with a domino effect on how pages are understood: the syntax and use cases of robots.txt are worth checking before writing a broad rule. One simple rule then closes the door on the most common fault: if an area is closed to crawling, avoid feeding it through internal linking. Otherwise you deliberately create crawl dead ends and dilute the logic of the navigation.
The rest is a matter of objective. Blocking crawling and blocking indexing do not answer the same intention, and confusing them produces the most frequent sequencing error: if you have to make an already known URL disappear, do not block the robot's access first — Google must be able to re-crawl in order to see the noindex, the 301 or the 410. For non-HTML content (PDFs, images), the X-Robots-Tag HTTP header is the only way to convey a noindex. And for a staging environment or a genuinely private space, protection by authentication (of the .htpasswd kind) blocks access completely, for robots and users alike: it is more reliable than a simple robots.txt, which stays public and does not prevent the indexing of already known URLs.
Concentrating crawling on what counts
Crawling is not infinite, and it concentrates on what you give it to see. Two movements steer it, in this order: remove what wastes it, then make what deserves revisiting easier to reach.
Spotting crawling that is badly spent
A badly used crawl budget is rarely spotted with a single metric. Look instead for converging signals:
- the recurring crawling of parameters and variants with no value (sorting, combined filters);
- a high volume of “discovered” URLs for few URLs actually indexed;
- spikes of server errors (5xx) or slowdowns correlated with a drop in crawling.
The third deserves keeping as it stands: an infrastructure incident can have a delayed SEO effect, through a slowdown in re-crawling and therefore in page updates. The best gains then come from removing the paths that create parasitic URLs: clean up internal linking so that it does not point to useless parameters, frame faceted navigation by leaving indexable only the combinations that answer a real search intent, prevent internal search from generating endlessly crawlable pages, and stabilize canonicalization on the variants (www or not, trailing slash, http or https, parameters). The idea is not to block everywhere: preventing discovery — no internal links, no presence in the sitemap — often stays cleaner than letting discovery happen and then trying to make up for it with directives. And when the volume itself becomes the problem, the trade-off between server capacity and crawl demand belongs to the SEO crawl budget.
Prioritizing high-impact pages
Effective crawling serves your business strategy: categories, conversion pages, pillar content, pages that already carry impressions and clicks. On high-volume sites — catalogues with thousands of products and hundreds of categories, a situation frequently seen in e-commerce — the point is to bring those pages up in the internal linking and to limit the variants. The orders of magnitude in the SERP set the frame: the click-through rate on the first organic position (desktop) is 34% (SEO.com, 2026), and the rate on page 2 of the SERPs is 0.78% (Ahrefs, 2025). Crawling is not an end in itself: it must serve access to the positions that count.
Running a crawl analysis: scope, cross-checking, decisions
A crawl quickly produces thousands of rows. What makes it an analysis is a scope decided in advance, a cross-check with what Google observes, and an output in decisions rather than in a list of anomalies.
Framing the crawl, then cross-checking what it shows
A useful crawl starts with a clear scope: which segment do you want to secure — blog, categories, products, local pages? Which objectives: uncover orphan pages, map the redirects, measure depth, identify duplication by parameters, check the rendering of a JavaScript template? On very large sites, favour an approach by batches (directories, page types, templates) rather than a page-by-page reading: it is the only way to identify root causes — a global canonical rule, a redirect pattern, a resource block — that affect hundreds or thousands of URLs.
An external crawl shows what the robot can explore; to know what Google does, cross-check with Search Console: indexing states, excluded pages, errors, sitemaps. A large volume of pages “crawled, currently not indexed” or “discovered, currently not indexed” serves as a warning, but the two statuses do not say the same thing: the first describes a page seen and then set aside, the second a URL known and not yet crawled. Its limit has to be known: it tells you what Google observes and decides, but not always why an architecture or an internal linking scheme generates so many parasitic URLs. Audience data, for its part, does not measure crawling; it prevents a frequent trap, fixing anomalies on pages that have neither traffic, nor conversion, nor anything at stake.
When an important page is almost never crawled
The diagnostic sequence comes in three stages. First check discoverability: does the page receive internal links from strong pages, is it too deep, is it missing from the sitemap or caught in a canonicalization conflict? Then check the technical brakes: latency, 5xx errors, redirect chains. Finally cross-check with Search Console: “discovered, currently not indexed” means that Google knows the URL but has not yet crawled it. Four stages are then distinguished — discovery, crawl scheduling, the capacity to carry it out, and only then the evaluation of the content — and thin content is only one hypothesis among others: queuing, unfavourable prioritization or server load explain the status just as well.
The finding then converts into actions, reasoned by template. The quick wins: fix an over-broad robots rule, remove internal links pointing to redirects, repair a 404 pattern on a template. The structural projects: rebuilding the facets, consolidating the canonicals, simplifying the pagination, improving the accessibility of the rendered content. The objective is not “zero alerts”, but technical stability that lets Google discover, render and process the important pages over time.
Tracking crawling over time
A one-off crawl photographs a state; it does not see what breaks “silently” between two releases. Yet that is exactly where the most expensive regressions sit: redirects appearing on a whole template, an explosion of parameter URLs, new orphan pages, blocked resources, or a rise in server errors. A regular check catches them before they turn into a loss of indexing.
The indicators to follow are indicators of stability and focus: fewer parasitic URLs discovered, fewer redirect chains and internal 4xx, fewer pages excluded for duplication, an improved “submitted” vs “indexed” gap on the sitemap, and rising impressions and clicks on the priority segments. What matters is to measure in batches — templates, directories, page types — rather than page by page.
That leaves connecting the three planes: crawling (the crawl findings), indexing (what Google decides) and value (what the pages bring in). It is that cross-reading that allows you to prioritize the fixes protecting pages that matter, rather than optimizing sections with no impact, and it is the most reliable way to measure the real effect of a crawl optimization over time, before and after, on comparable segments. One task resists doing by hand: comparing two successive crawl maps to spot what changed between two releases — new redirects, parameter URLs that have appeared, pages that have dropped out of the internal linking. That is what the audit and mapping module covers.
FAQ on site crawling and SEO crawling
How do you launch a site crawler step by step, without skewing the results?
Define the scope (domain, subdomains, directories), then set a URL ceiling if the site is large. Start from representative entry URLs — home page, hubs, categories — and check that the crawler follows crawlable HTML links without getting lost in parameters. Finally export the URL lists by status, depth and directive, so as to identify patterns rather than isolated cases.
How does Google crawling work, in practice?
Google discovers URLs through links, internal and external, and through sitemaps, then places those URLs in a queue. Googlebot then fetches the content and, depending on the case, the resources needed for rendering. Nothing is instant: the trade-off depends on perceived quality, authority, freshness and server constraints. The sitemap flags pages added or modified, without guaranteeing an immediate visit.
What is the difference between crawling and indexing?
Crawling corresponds to visiting a URL and fetching its resources. Indexing corresponds to the decision to add that URL, or its canonical version, to the index so that it can appear in the results. A page can be crawled without being indexed — noindex, duplication, insufficient quality. Without indexing, it cannot rank in the SERPs.
Which tools should you use to crawl a website without multiplying solutions?
To understand what Google sees and decides, rely on Search Console: it exposes indexing coverage, exclusions, errors and sitemaps. To map the site “like a robot” and detect technical patterns — redirects, depth, canonicals, orphan pages — use a diagnostic crawler. The two are complementary: one says what Google does, the other why.
Does a sitemap guarantee crawling and indexing?
No. A sitemap serves to report pages added or modified, but does not force crawling. And even once crawled, a URL may not be indexed — noindex, duplication, low value. The sitemap becomes genuinely effective when it is “clean”: URLs returning 200, indexable, canonical, and aligned with the internal linking.
Why does Googlebot crawl useless URLs (parameters, filters) and how do you avoid it?
Because those URLs exist and are discovered: internal links from filters and sorting, pagination, internal search, external links, or sitemaps that are too permissive. To avoid it, reduce rediscovery — remove internal links to non-strategic variants, clean the sitemap, stabilize the canonicals, frame faceted navigation. Blocking can help, but it must not become a sticking plaster for an architecture that produces too many URLs.
What should you do if important pages are almost never crawled?
Go back to the three stages of the diagnosis: discoverability first — internal links from strong pages, depth, presence in the sitemap, canonicalization conflicts; technical brakes next — latency, 5xx, redirect chains; cross-check with Search Console last. If the pages appear there as “discovered, currently not indexed”, Google knows the URL without having crawled it yet: examine the queuing and the crawl capacity — internal linking, depth, latency — before concluding that the content is thin.
Do external links influence discovery and crawl frequency?
Yes. Backlinks make URLs easier to discover and can reinforce their perceived importance, which plays into prioritization. In practice, an external link to an orphan page can keep it reachable, but it does not replace clean internal linking if you want stable crawling over time.
Robots.txt or noindex: which to choose, depending on the objective?
Choose according to the expected outcome:
- Prevent indexing: noindex in a meta robots tag, or X-Robots-Tag for non-HTML content;
- Limit crawling: robots.txt, to avoid wasting crawling on worthless areas;
- Block access completely (confidential material, staging): authentication.
If a URL is already known and you want it to disappear, avoid blocking crawling too early: Google has to re-crawl in order to see the signal.
Which indicators should you follow to measure a crawl optimization over time?
Follow indicators of stability and focus: fewer parasitic URLs discovered, fewer redirect chains and internal 4xx, fewer pages excluded for duplication, an improved “submitted” vs “indexed” gap, rising impressions and clicks on the priority segments. Measure in batches rather than page by page.
What are the SEO risks with a site heavily dependent on JavaScript?
Rendering costs the engine more, which can delay or prevent indexing. Check what is really present in the rendered HTML, the discoverability of internal links — a link that appears only after a script has run is not a reliable crawl path — and access to the content without depending on complex scripts.
Continue reading
- The indexing reports are no longer enough and you need the trace of what Googlebot actually requested, URL by URL: the log analysis shows the visit frequency per section, the statuses actually returned and the crawl waste.
- You have to check what Google does with your URLs and are looking for where to read coverage, exclusions and errors: the Google Search Console reports and how to interpret them, report by report.
.png)
%2520-%2520blue.jpeg)

.jpeg)
.jpeg)
.avif)