Tech for Retail 2025 Workshop: From SEO to GEO – Gaining Visibility in the Era of Generative Engines

Back to blog

Technical SEO audit: what blocks access, and in what order to fix it

SEO

Discover Incremys

The 360° Next Gen SEO Platform

Request a demo
Last updated on

26/9/2026

Chapter 01

Example H2
Example H3
Example H4
Example H5
Example H6

Published pages that never show up, rankings that move after a release, a tool report thousands of lines long: these situations raise a question of access first.

 

What a technical audit checks, and how you sort what it surfaces

 

A technical audit checks what determines the ability of search engines to crawl, render and index your pages. It replaces neither content analysis, nor authority analysis, nor business analysis — even though technical findings have to connect back to them. The overall scope of an SEO audit sets the frame; at URL level, this layer prevents three very costly scenarios:

  • The page is not discovered (no internal links, excessive depth, inaccessible pagination, JS that hides the links);
  • The page is discovered but poorly understood (duplicates, inconsistent canonical, competing versions, contradictory markup);
  • The page is understood but “too expensive” to process (slowness, server instability, heavy JavaScript rendering), which delays indexing.

The opportunity cost is immediate: the click-through rate on the first organic position (desktop) is 34% (SEO.com, 2026), and the rate on page 2 of the SERPs is 0.78% (Ahrefs, 2025) — benchmarks collected in the SEO statistics. The right objective is not “zero alerts”, but technical stability that lets Google discover, render and process your important pages, and lets your users reach them fast and without friction. Since Google ships 500 to 600 algorithm updates a year (SEO.com, 2026), continuous improvement is worth more than a frozen state.

 

The external crawl: what it shows, what it does not tell you

 

A technical audit relies on an external crawl: you observe the site as a robot does, without depending on the CMS, because the analysis bears on visible symptoms — URLs, internal links, statuses, directives, rendered HTML, canonicals, depth. It reveals what the eye does not see: redirect chains, deep pages, non-indexable URLs linked internally, inaccessible pagination. Its limit decides what comes next: a crawl shows what the robot can explore, not what Google does. Hence the mandatory cross-check with Search Console — indexed pages, excluded pages and exclusion reasons, errors. These reports have their own latency and their own sampling: a fault established by a crawl or a URL inspection is real before it appears there, and the absence of a trace is not a refutation.

 

Blockers, amplifiers, marginal optimizations

 

Raw audits confuse exhaustiveness with usefulness: you produce a thick report without the teams knowing what to do on Monday morning, when 10 prioritized decisions are worth more than a backlog of 500 unsorted tickets. Classify at three levels:

  • Blockers: they prevent the crawling, rendering or indexing of the pages that matter (over-restrictive robots.txt, polluted sitemap, redirect loops, dead-end pagination, JS that hides the links).
  • Amplifiers: they improve consolidation and readability (consistent structured data, fewer chains, better accessibility of deep pages).
  • Marginal optimizations: useful but secondary as long as the fundamentals are not stable.

Within each level, sort by impact, effort and risk of regression, and apply the grid in batches — templates, directories, business segments — rather than page by page: a global redirect rule comes before a handful of missing titles when the objective is to secure indexing.

 

Crawling and indexing: access directives, sitemap and crawl budget

 

Crawling is not unlimited: worthless URLs, redirects and duplicates consume it at the expense of strategic pages. Two files decide most of it — the one that says where the robot need not go, and the one that says where it should go first.

 

Robots.txt: the three checks and the safeguard

 

The file lives at the root and steers crawling. A useful gatekeeper, it becomes dangerous when badly set: an over-broad directive can forbid a whole directory, or the whole site. Check three things: that strategic directories are not caught by an over-broad rule, that the resources needed for rendering (CSS, JavaScript, critical images) are not blocked, and that the file in production is not a copy of the staging one — which happens at every migration.

The safeguard fits in one sentence: if an area is closed to crawling, it should not be fed by internal linking. Otherwise you create crawl dead ends. And one distinction avoids the costliest confusion on the subject: blocking in robots.txt is not deindexing — a URL closed to crawling can remain indexed without its content, whereas a noindex only takes effect if the page remains crawlable, since the robot has to be able to read the directive to apply it; a page blocked by robots.txt will therefore never leave the index through a noindex it will not go and fetch. Nor does blocking an area settle the reason it was generated in the first place: robots.txt is not a sticking plaster for an inconsistent architecture.

 

Sitemap and volume: the gap between submitted URLs and indexed URLs

 

A sitemap lists the URLs intended for indexing, and its value comes from its cleanliness: it reflects your indexing strategy, not the inventory of existing URLs. Leave in it only URLs that answer 200, indexable and canonical. Then watch it: the “submitted vs indexed” gap is often more instructive than a plain “OK” status. A high volume of pages “crawled, currently not indexed” signals a problem of perceived quality, duplication or architecture, not a fault in the file.

What the robot discovers, and how often it comes back, belongs to SEO crawling. Past a certain volume, the problem changes in nature: it is no longer about being crawled but about being crawled usefully — the trade-off between server capacity and crawl demand then belongs to the SEO crawl budget.

 

Architecture, depth and orphan pages

 

The deeper a page sits, the harder it is to discover and the less internal weight reaches it. A practical rule often used is to aim for important pages reachable in around three clicks, through topical hubs and contextual links. Three points are checked: crawlable links, present in the HTML and reachable without complex interaction; a consistent distribution, with the pages that carry the business receiving more links; explicit handling of the non-indexable areas — basket, account and checkout must remain reachable by users, it is their indexability that gets settled, not their presence in the internal linking. What is at stake here is discoverability, not editorial strategy.

An orphan page no longer has an internal link from the rest of the site. The treatment depends on its value: a useful page, attach it to a hub (category, topic page, parent page) and create contextual links; an obsolete page, remove it cleanly (410 or 404 depending on strategy) or redirect it if an equivalent exists; a low-value but necessary page, such as legal notices, do not over-link it needlessly.

One clarification avoids the most common misreading: “orphan” describes the absence of internal links, not the absence of links altogether. The page can receive inbound links, be visited and even generate revenue while remaining isolated from the robot's point of view. The case is classic after a migration: the internal linking has gone, the inbound link profile remains, and re-crawling degrades without traffic announcing it.

 

URLs, redirects, canonicals and pagination: one version, one path

 

A large share of technical problems comes from URL conflicts: several paths lead to the same content, or one page replaces another without the signals consolidating. The engine then chooses in your place, and it chooses badly.

 

Redirects: chains, loops and internal links that point elsewhere

 

A 301 signals a permanent move: migration, URL change, http to https consolidation. A 302 announces a temporary state: kept in place for a long time on an indexable page, it generally ends up being treated as permanent, so that the problem is not a mechanical loss of signals but the ambiguity of what is declared — the stated intention no longer matches the situation, and the moment of the switch remains in the search engine's hands. Two rules are enough: shorten, a redirect must be direct (A→B), never A→B→C; fix the internal links, because internal linking that points to redirected URLs slows down crawling and rendering. Loops (A→B→A) break crawling. A redirect that affects a whole template comes before an isolated anomaly.

 

Canonicals and competing versions: the four-point alignment

 

A canonical becomes dangerous as soon as it contradicts reality: pointing to a non-indexable page, applied globally to the home page, or inconsistent with an active redirect. Duplicates almost always come from competing versions — http and https, www and non-www, trailing slash or not, sorting, pagination and tracking parameters. Treat duplication as a matter of system consistency: not a series of micro-fixes, but an alignment between (1) the version served, (2) the version linked internally, (3) the version listed in the sitemap, and (4) the indexed version observed in Search Console. Consolidating on HTTPS, with no insecure resource, is part of it. On a multilingual site, two further checks apply: the reciprocity of language declarations and their consistency with the canonical.

 

Pagination and infinite scroll: reaching page 3

 

Pagination must let the robot reach deep products or articles. Common faults: AJAX pagination that cannot be detected, noindex on paginated pages, blocking in robots.txt, a systematic canonical to page 1. The last one deserves naming: if every paginated page canonicalizes to page 1, you are sending a contradictory signal — “pages 2 and beyond exist, but do not count”. The engine crawls less and you lose the items that only appear on page 3. The configuration that holds is a self-referencing canonical on each page, with crawlable pagination links.

Infinite scroll carries the same risk: if the additional content does not exist through reachable URLs, it becomes hard to crawl and to index. Three remedies — provide a paginated alternative or dedicated URLs, make sure internal links allow discovery beyond the first screen, and check the rendering so that what appears to the user is reachable by the engine.

 

HTTP statuses: what the server really answers

 

The useful sorting separates two cases that are often confused. Internal 404s — links from the site to a URL that does not exist — come first, because they create crawl dead ends; fix them at source (template, menus, “related products” modules) rather than through a blanket redirect. External 404s are decided case by case: redirecting them all to the home page is rarely a good idea. For soft 404s, the principle is binary: either the page does not exist and returns a real 404 or 410, or it exists and offers real, indexable content.

5xx errors follow another logic: an availability incident with an SEO effect. Three questions frame it — quantify the frequency and the scope, check whether the errors coincide with traffic peaks, batch jobs, deployments or external dependencies, and check the effect on indexing. This work comes early, because it conditions everything else.

Finding Effect on the engine side Expected decision Acceptance criterion
404 on a URL linked internally Crawl dead end Fix the link in its template No internal link to a non-200 URL
404 on a URL linked from outside Benefit of the inbound link lost Redirect to an equivalent, otherwise 410 No mass redirect to the home page
200 on a page with no content Soft 404: crawling wasted Decide: a real 404, or indexable content Status consistent with what is displayed
Recurring 5xx on a key template Unstable indexing Handle as an incident, before anything else Errors back to the baseline level

 

Performance and JavaScript rendering: the processing cost of a page

 

A page has a cost: fetching it, executing it, understanding it. That cost decides how quickly your pages enter the index.

 

When slowness becomes an indexing problem

 

A heavy, unstable page is expensive to process: the engine comes back to it less often and indexing falls behind. The signal to watch is therefore not a score, but the conjunction of a heavy template and pages discovered but not indexed on that template. On the user side, the cost is immediate: the mobile abandonment rate when loading exceeds 3 seconds reaches 53% (Google, 2025). A quick win remains useful — spot images above 500 KB and compress them, a practitioner's working threshold rather than a standard. The measurement itself and the trade-off between cost and gain belong to the website performance audit.

 

JavaScript rendering: what the robot actually gets

 

JavaScript is not bad by nature. The risk appears when the content, the internal links or the metadata depend on complex or fragile rendering: the engine then indexes incomplete content, delays indexing, or misses links and therefore whole pages. The question to ask is not “is there JavaScript?” but “are the essential content and links present in the rendered HTML, stably and quickly?” The check is concrete: compare the raw HTML with the rendered HTML; if links or blocks appear only in the second, you have a rendering dependency — single-page applications and filter pages first.

 

Extractability: structured data and the readability of rendered HTML

 

A page is also read by answer systems that extract fragments from it: the more readable and consistent it is on the HTML side, the simpler it is to cite. A technical audit does not create that citability; it removes the obstacles to it.

Four types of markup recur, each with its condition of use: Article for editorial content, provided the declared properties (date, author, main image) really exist on the page; FAQPage when the page carries a visible FAQ; HowTo if the page describes a real procedure, in steps; BreadcrumbList to reflect the architecture. Two checks come first: validity — no syntax errors and no missing required fields in the JSON-LD — and alignment — the markup must reflect the visible content, same labels, same questions, same items. The two most frequent faults: a generic template that injects inaccurate properties (missing author, wrong date), and a marked-up FAQ whose questions are not displayed. Hence a maintenance rule: if a piece of data changes in the content, it must be modified in the same place — or in the same flow — as its structured equivalent, which limits the debt at redesign time.

That leaves the readability of the HTML: a clear, non-contradictory heading hierarchy, and summary formats where they bring real value. Two warning signals show up in the crawl: very little visible content relative to the complexity of the rendering, which can indicate partial indexing, and hidden content, through CSS for instance, not intended for the user but injected to add volume. The principle: what has to be understood and cited must be reachable, visible and stable in the rendering.

 

From finding to ticket: recommendation, acceptance testing and non-regression

 

A technical audit is a decision, not a document: it turns into an executable backlog, then into checks after deployment. With one reservation — the effects are rarely visible within a few days: expect gradual signals.

 

Writing a recommendation a developer can take as a ticket

 

Impose a repeatable format, in four fields: hypothesis, why this point may hold visibility back; evidence, a crawl extract, a Search Console report, an analytics segment; fix, a redirect rule, a template correction, a robots.txt or sitemap adjustment, a rendering change; acceptance criteria, what must be true after deployment. Group the actions by template and by directory. Avoid vague wording — “improve speed” — and favour testable actions: “no redirect chain left on /x/”, “every sitemap URL answers 200”.

 

The deployment sequence: the order follows the blocker observed

 

The order in which the fixes ship is not neutral, but it is deduced from the blocker observed rather than from a list valid everywhere: what prevents access comes before what improves understanding. (1) Secure crawling: robots.txt and sitemap. (2) Stabilize the URLs: HTTP statuses, direct redirects, removal of chains and loops. (3) Make the content reachable: JavaScript rendering and paginated journeys, a priority as soon as blocks or links exist only in the rendered HTML — content inaccessible to rendering comes before any markup. (4) Make understanding reliable: consistent structured data. (5) Reduce the processing cost: dependencies and weight. On a site with no dependence on JavaScript, step 3 reduces to pagination and markup moves up accordingly.

The workload depends on five variables, and it is on those that a budget is framed: site size, number of templates, migration history, dependence on JavaScript, and the expected level of reporting — a plain report or a backlog ready to develop. On a fixed fee, on a time-and-materials basis or per project, it is these variables that move, not the prioritization rule.

 

Validating the fixes, and knowing when to run the check again

 

After each release, validation comes in three steps: a re-crawl targeted on the modified directories, a check in Search Console (indexing, exclusions, errors), then the measurement of the effect on the pages affected. The crawl shows what Google can explore; Search Console helps you understand what Google actually does. On cadence: at least one structured technical review a year, and after every redesign, migration or major update; then continuous monitoring if the site moves fast — frequent deployments, growing volume — to catch regressions before rankings.

Two signals finally say that the limitation is no longer technical: well-optimized, well-indexed pages that plateau despite a mastered intent, and a lasting gap between your perceived quality and your visibility on competitive queries. Grouping the findings of an external crawl and the indexing states from Search Console into a single reading, sorted by template, is the task covered by the audit and mapping module.

 

FAQ: understanding the technical side of an SEO audit

 

What is a technical SEO audit?

 

It is a structured analysis of the elements of a site that determine the ability of search engines to crawl, render, understand and index the pages: access and noindex directives, sitemap, HTTP statuses, redirects, canonicals, depth and internal linking, JavaScript rendering, structured markup. It replaces neither content analysis, nor authority analysis, nor business analysis — it conditions their effectiveness.

 

How do you run a technical SEO audit, step by step?

 

First map the site with an external crawl: URLs, links, statuses, directives, canonicals, depth. Then cross-check with Search Console to separate what the engine can do from what it does. Classify the findings into blockers, amplifiers and marginal optimizations, then sort them by impact, effort and risk, applied in batches. Finally deploy in order, re-crawl and check indexing.

 

Which technical factors have the most impact on organic search?

 

In practice, the list is short: indexability and access directives, HTTP statuses, consistency of URL versions, handling of duplicates through canonicalization, depth and linking towards business pages, and the accessibility of the rendering when JavaScript drives display and links. The rest — micro tag anomalies, non-blocking warnings — comes afterwards, unless a demonstrated impact is found on a family of URLs that matter.

 

How do you prioritize technical problems when the list is too long?

 

Group by families — templates, directories, business segments — then sort by the risk of preventing crawling or indexing, the volume of URLs concerned, the presence on pages that matter, and finally the effort and the risk of regression. Do not aim for “zero warnings” but for the maximum measurable impact: ten well-prioritized decisions are worth more than a backlog of five hundred unsorted tickets.

 

Why is an external crawl essential in an SEO-oriented audit?

 

Because it observes the site as a robot does, without depending on the CMS or the stack: URLs, internal links, statuses, directives, rendered HTML, canonicals, depth. It is the most reliable way to map what exists and to spot traps that are invisible from the admin interface — redirect chains and loops, deep pages, non-indexable URLs linked internally, inaccessible pagination.

 

Which technical problems are the priority “blockers” in an audit?

 

In order: access blockers (over-restrictive robots.txt, unintended noindex, strategic pages out of reach), recurring server errors, lost or ambiguous signals (404s on useful pages, redirect chains, long-lived 302s), then duplication conflicts (URL versions, canonicals, parameters). Depth, linking and finishing optimizations come next, especially where they affect key templates.

 

What is an inconsistent canonical and why is it risky?

 

A canonical becomes inconsistent when it contradicts the technical reality or the indexing strategy: it designates a non-indexable page, it points globally to the home page, or it does not match the version actually served. The risk is twofold: accidentally de-indexing useful pages, or diluting the signals across competing versions instead of consolidating them on a single one.

 

How can an orphan page have links?

 

“Orphan” describes the absence of an internal link path from the rest of the site, not the absence of links altogether. The page can receive external inbound links, appear in the sitemap, or belong to an island of pages that link to each other without being connected to the main site. It therefore remains discoverable and can generate revenue while staying fragile: without internal linking, its re-crawling and the consolidation of its signals degrade.

 

Why should paginated pages keep their own canonical?

 

Because pages 2, 3 and 4 carry different listing content and give access to the depth of the catalogue. Canonicalizing them all to page 1 amounts to saying that they exist without counting: crawling shrinks, deep content becomes harder to discover, and the internal linking contradicts the canonical. A self-referencing canonical on each page of the series preserves crawlability and consistency.

 

Should the technical SEO audit evolve with GEO (Generative Engine Optimization)?

 

Yes, but by extension rather than by substitution. Crawl and indexing checks are joined by extractability checks: relevant and valid structured markup, a consistent heading hierarchy, essential content really present in the rendered HTML, summary formats on the pages that answer questions. The aim is not to replace SEO, but to remove the obstacles that keep a page from being read and picked up.

 

What should be checked on an international site (hreflang)?

 

The audit checks the consistency of the language and country pairs, the reciprocity of the annotations, and the alignment between hreflang, canonicals and URL structure. Without that consistency, you risk impressions “in the wrong place” and a dilution of the signals across versions.

 

How do you handle duplication and canonicalization in an SEO audit?

 

Technical duplication — http/https, www/non-www, trailing slash, parameters — and content duplication — pages that are too similar, facets, badly handled pagination — create competing URLs. The audit checks that a single canonical version exists per piece of content, and that the canonical tags remain consistent with the redirects and with the declared indexability.

 

Continue reading

 

  • The crawl and Search Console contradict each other, and you need proof of what the robots actually requested: the log analysis shows the visit frequency per section, the statuses actually returned and the crawl waste.
  • You have to write or fix access rules and want the detail of the directives: the syntax and use cases of the robots.txt file, with what it does and what it does not do.
  • Your site serves several languages or several countries and targeting becomes the subject: international SEO covers language and geographical targeting, language declarations included.

Discover other items

See all

Next-Gen GEO/SEO starts here

Complete the form so we can contact you.

The new generation of SEO
is on!

Thank you for your request, we will get back to you as soon as possible.

Oops! Something went wrong while submitting the form.