Tech for Retail 2025 Workshop: From SEO to GEO – Gaining Visibility in the Era of Generative Engines

Back to blog

Claude AI Agent: Claude Code, Sub-Agents and Controlled Delegation

GEO

Discover Incremys

The 360° Next Gen SEO Platform

Request a demo
Last updated on

26/9/2026

Chapter 01

Example H2
Example H3
Example H4
Example H5
Example H6

What sets a Claude agent apart: long context and delegation

 

An assistant hands back text; an agent acts on an environment, observes what it has produced and starts again. That switch changes everything: as soon as the system writes into files and runs commands, what you are deciding is no longer the quality of an answer, it is a scope of action. On Claude agents, two properties govern that scope: the amount of context the agent holds in one go, and its ability to hand part of the work to specialized sub-agents. Permissions, budgets, logging: the rest follows from those.

The maturity of the tool is no longer the question: 70% of Fortune 100 companies are equipped (Thunderbit, 2025). If the trade-off between families of tools is not yet settled at your company, the switching criteria and the requirements to set before signing are gathered on AI agent platforms. Here the tool is assumed to be chosen: what follows is the scoping of a Claude agent on a real code base.

 

Long context is a resource to arbitrate, not a selling point

 

A large context window does not make an agent better: it gives you room to manoeuvre, and it is up to you to decide how to spend it. The trade-off comes up on the first slightly broad task, and it has only two outcomes. Either you load everything into a single context: the agent sees the whole scope, nothing is lost between steps — but the window fills up, and the last steps reason on a saturated context where the initial instruction no longer counts for much. Or you split: each piece restarts on a clean context and reasoning quality stays stable — but you have to prime it again every time, and that priming has a price.

That choice determines your agent’s architecture, and it is made before the first run, not when results start to degrade. Two symptoms signal a saturated context: the agent asks again for information it has already been given, or it contradicts a decision it made itself twenty steps earlier. When either appears, the task is too broad for a single context.

 

A long run changes what you have to watch

 

The second property comes down to duration: a Claude agent can carry a long task through without the hand being taken back from it. The autonomy durations that circulate mean nothing until they are tied to a model, a version and the protocol that measured them — remember the property, not the ceiling. And that property is not a promise of results, it is a supervision constraint. An agent that sustains a long run is an agent nobody watches while it works, and that moves control entirely: you are no longer watching a screen, you are rereading a trace after the fact.

Three consequences follow, and they structure everything that comes next. You need a budget, without which nothing stops an agent going round in circles; a usable log, without which you will not know what happened during the hours you were not there; an explicit stop criterion, without which the agent will decide for itself that it has finished. Those usage and performance benchmarks, with their sources, are gathered in our set of Claude statistics: they give an order of magnitude, not your results.

 

Claude Code: an agent that acts on the project, not in a window

 

The difference between a chat and a development agent is not a difference of model, it is a difference of interface — and it changes everything. In a conversation window, you describe your problem, you receive code, you copy it over, you run it yourself and you come back to report the error: the round trip is manual, and you are the loop. With Claude Code, the agent operates from the terminal, directly on the project: it reads files, modifies them, creates them, runs commands, launches the tests, reads the errors and starts again. You no longer copy anything.

That shift is what turns an advisory assistant into an agent that runs a full cycle, and the tool is widely present in teams: developers using Claude or Claude Code are estimated at between 41 and 68% (Faros AI, 2026), a wide range to be read as an order of magnitude. It is also what makes the question of scope urgent: an agent that only acts in a window can break nothing, a Claude Code agent that acts on the project can.

 

The plan, execution, observation, correction loop

 

A reliable development agent follows an explicit iterative loop, and that loop is what you frame, not the code it writes. To make it controllable, impose a standard output at each step: without it, you get a final result with no idea how it was obtained, which amounts to rereading the whole repository.

  • Plan: steps, files affected, assumptions, stop criteria.
  • Execution: list of actions (creations, modifications), commands run.
  • Observation: test results, useful logs, errors encountered.
  • Correction: fixes applied, justification, fresh test pass.

The step most often neglected is the first. A plan handed over before any modification gives you a free stopping point: you see which files are going to move, you spot a scope that is too broad, and you correct the instruction before the agent touches the repository. It is the only moment in the cycle when an error costs nothing.

 

What the loop really produces on a repository

 

The tasks this loop holds up on are the ones whose deliverable can be checked mechanically. Refactoring or migrating is not “changing code”, it is preserving a behaviour: require a proof strategy — tests before and after, a readable diff, incremental steps. Writing tests is the best lever for securing an autonomous agent, because the success criterion is verifiable without human review: the agent writes targeted tests, runs them, then fixes until green.

Documentation obeys a rule of its own, because nothing in it fails visibly: a false text reads as well as a correct one. To avoid purely plausible documentation, force the agent to cite only things it has actually seen or run: file paths, function signatures, commands executed and their output. A format that holds: prerequisites, installation, commands, troubleshooting for common problems, reproducible examples, and a “what the module does” section followed by “what it does not do” — that last line prevents more downstream usage errors than any other.

 

Sub-agents: splitting by deliverable, and paying for the split

 

Delegating to specialized sub-agents becomes relevant when a task genuinely parallelizes: front end, back end, tests, review. The trigger is not the size of the job, it is the presence of independent pieces, each judged on its own deliverable without knowing the detail of the others. If two pieces have to wait for each other at every step, splitting them saves no time: it adds boundaries, and therefore chances to lose information.

The second trigger is context: a task that saturates the window gets split even when it is sequential, because each sub-agent then restarts on a clean context. Beyond three or four sub-agents, you are no longer scoping a delegation but a coordination, with its priorities, its conflicts and its restarts: the patterns that structure it belong to AI agent orchestration.

 

Naming each sub-agent after its deliverable

 

To avoid losing context, structure your sub-agents by deliverable rather than by vague intent. A sub-agent “that improves quality” hands back nothing verifiable and consumes budget until you stop it; a sub-agent defined by what it has to produce is assessed in a minute. Three splits work almost every time:

  • tests sub-agent: adds or updates the test suite, then runs the test command.
  • migration sub-agent: applies a plan file by file and supplies a diff.
  • review sub-agent: lists risks, antipatterns and security impacts.

The rule is checked in only one way: if you cannot say what the sub-agent has to hand back or how you will know it has finished, it is not ready to launch. Write its deliverable before its prompt. And give it an explicit file scope: a “review” sub-agent that modifies code has crossed its boundary, even if the change is a good one.

 

What passes between two sub-agents, and what the split costs

 

The weak point of a delegation is not the work of each sub-agent, it is the handover from one to the next. Never let a sub-agent pass its raw context to the following one: impose a short, standardized summary, always the same — what has been done, what is still open, the decisions taken and their one-line justification, the files touched. That fixed format lets you reconstruct the sequence without reopening every run, and it stops the next sub-agent reinterpreting a context it did not live through.

The split has a price, and it has to be stated plainly: each sub-agent consumes its own budget. It restarts from an empty context, which has to be primed again with the instructions, the repository conventions and the state of progress — and that priming is paid for as many times as you split. That is the trade-off to which a context window gives its unit — provided it is tied to a dated model, for example 200,000 tokens for Claude Opus 4.5 (llm-stats.com, 2026), the value changing from one version to the next: as long as the task fits comfortably inside it, a single agent costs less and loses less information; the moment it overflows, the priming overhead becomes the price of reasoning quality at the end of the job.

 

The output contract: schemas, evidence and stop criteria

 

To make an agent reliable, think “contract”. On the input side you supply a context, constraints and an objective; on the output side you require a stable schema that makes automation possible — programmatic reading, routing rules, storage. A useful agent does not just answer: it produces auditable artefacts — diffs, logs, test results, classifications, decision lists — and review is done on those, never on a prose summary.

 

Five mandatory fields, and a prompt that steers a process

 

The output contract for a refactoring task fits in five fields, and all of them are requested:

  • files_changed: list of paths and type of action (creation, modification, deletion).
  • diff_summary: a summary in 5 points maximum.
  • commands_run: commands and results.
  • risks: points to review manually.
  • stop_reason: finished, to be reviewed, or blocked.

The last field is the most useful of the five: it makes a run sortable without being read. A finished run goes to normal review, a run to be reviewed goes up into a human queue, a blocked run calls for a restart; you sort fifty runs in a few minutes instead of opening them all. All of this assumes an instruction written accordingly: the prompt has to steer a process, not an answer. Four elements are therefore mandatory on every task — an output format (tagged structure, sections, checklist), expected evidence (diff, test command and its result, files modified), a stop criterion (tests green, minimum coverage, or escalation) and a budget (number of iterations, maximum time, maximum cost).

 

Sources of truth and debugging: escalate rather than invent

 

A generative model remains probabilistic: it produces the most plausible continuation given its context, with no guaranteed understanding. The consequence stands as a design rule: if you do not feed your agent sources of truth, it will compensate with the plausible. So impose an explicit hierarchy, from the most reliable to the least:

  • the repository’s code and tests, as operational truth;
  • versioned internal documentation (architecture decisions, technical notes, getting-started files);
  • tickets and merge requests, as a history of decisions;
  • failing that, human escalation instead of invention.

The same requirement governs debugging, which almost always fails for the same reason: the agent proposes a fix before having reproduced the problem. Impose the full sequence — reproduce (commands, input and output), hypotheses ranked by likelihood, a minimal fix with an explanation, and locking it in with a test added or updated that fails before and passes after. That last condition is the one acceptance criterion that is not up for debate. Hence the general rule: do not ask for “correctness”, ask for “provability” — file paths, commands executed, test outputs, or escalation if the agent can cite nothing. Those deliverables must stay readable by your French-speaking teams; if your constraint is rather where the processing takes place, that is another subject, covered with the open models of the Mistral AI agent.

 

Permissions, budgets and loops: what stops the agent drifting

 

A Claude Code agent handles files, runs commands and automates version control actions. Faced with that, the question is not “can it do this?” but “within what scope?”. The answer is written once and holds for every run: it is the permissions policy. Here is the minimum scoping for a company, with the evidence to require in each case, because an authorization with no trace is not a controlled authorization.

Scope Allow by default Require human approval Evidence to require in the run
Reading files Yes, secrets excepted Access to sensitive folders List of the files consulted
Writing files On dedicated branches Security, auth and payment files Full diff and a 5-point summary
Running commands Tests, lint, build in a sandbox Deployment scripts, deletion, commands outside the sandbox Commands and outputs, secrets masked
Version control actions Local commit Push, PR to protected branches Target branch and run author

 

That grid is read line by line with the technical team, and each cell is discussed once only. What makes it workable is that it depends on no particular task: it holds for the first agent as for the tenth.

 

Escalation and budgets: stopping the correction loop spinning

 

In a company, the human in the loop is not a brake, it is an accelerator of reliability: the agent does the heavy lifting, the human approves the critical points. The approval trigger still has to be written down. A simple escalation model fits in three lines: full autonomy on reversible tasks (formatting, documentation, tests); mandatory approval on security, authentication, payment and dependencies; escalation if the tests fail after N iterations or in case of functional ambiguity.

That leaves the risk specific to agents that iterate: an agent can go round in circles — fixing one test while breaking another, rewriting without stabilizing. Prevention comes through budgets and non-negotiable success criteria:

  • an iteration limit (three correction cycles at most, for instance);
  • a scope limit (authorized directories and modules);
  • a cost and duration limit, settled before launch;
  • a mandatory final check (tests, plus a summary of the changes).

 

Secrets, isolated environment and clean failures

 

An agent that runs commands sees what your environment sees. And no command is safe by nature: running the tests, the lint or the build means executing code that comes from the repository — installation scripts, plugins, hooks, third-party dependencies — with the rights of the session. They are allowed by default because they are routine and reversible, not because they would be harmless: they call for the same isolation as the others. So isolate what it can see and do, to prevent unintentional exfiltration, mishandling of secrets or a destructive action: secrets outside the workspace (restricted environment variables, a vault), an isolated test environment, log masking (tokens, API keys, customer data), protected production branches with a mandatory merge request. Those four points are checked before the first run, not after the first incident.

Finally, an agent that acts has to know how to fail cleanly: the minimum is to avoid irreversible partial actions and dangerous repetitions. Four mechanisms cover most of the situations met in production.

Risk Mechanism Intended effect What stays with the human
Stalling on a long task Timeout and step-by-step resumption Avoid endless runs Set the threshold and the resumption step
Network or service instability Bounded retries with back-off Robustness without loops Decide the maximum number of retries
Double execution Idempotency through a run key No side effects Define what identifies a run
Regression introduced Rollback plus tests Fast return to a healthy state Choose between a fix and a rollback

 

Measuring an agent on a repository: logs, reproducibility, indicators

 

If your agent fails, you must be able to answer three questions: what did it do, why, and with what data? Without that, every incident turns back into a manual investigation. Logging is not an operations option, it is the condition for replaying a run and understanding a discrepancy. Five elements are logged systematically:

  • the system prompt, the user prompt and the parameters (model, temperature, limits);
  • the list of files consulted and modified;
  • the commands executed and their outputs, with secrets masked;
  • the results of tests, lint and build;
  • the run identifier, the timestamp, the duration and the estimated cost.

Then comes measurement, and this is where most teams pick the wrong unit: they look at the quality of the code produced rather than the cost of reviewing it. Measure on representative tasks, with a constant protocol, and track six indicators that hold together: the size of the diffs and the number of files modified; the test pass rate before and after; the number of iterations needed to stabilize; the human review time, in minutes, against agent time; the stabilization time, from the first failure to green; and the rate of regressions detected after merge.

The fourth is the one that decides, provided it is read correctly: what counts is not the ratio between agent time and review time, it is the sum of the two compared with the time the same work would have taken by hand. Ten minutes of agent followed by forty minutes of review makes fifty minutes: set against several hours of manual work, the operation remains highly profitable, even if the review weighs four times the machine time. The ratio only becomes an alarm signal when that total approaches the manual benchmark — and it is that comparison which answers the workload question when the time comes to extend use to other teams. The sixth measures what the agent really cost: a regression detected after merge cancels out several successful runs.

One last benchmark deserves to be put back in its place. A reference score is only to be cited tied to the model, the version and the evaluation protocol that produced it; taken out of that triplet, it compares with nothing and commits you to nothing. Even correctly attributed, it situates a level of general capability and says nothing about your repository, your conventions or your test coverage. A reference score is not an acceptance criterion: yours is written in green tests on your code base, and it is the only one that commits anybody.

 

FAQ on AI agents with Claude

 

What is a Claude agent?

 

A Claude agent is a system that uses the Claude models to plan and carry out tasks with a degree of autonomy, relying on tools: files, commands, integrations. What sets it apart from an assistant is not the quality of its answers, but the fact that it acts on an environment, observes the result of its action and starts again until a stop criterion you have set.

 

What are the advantages of Claude for agents?

 

Two properties really weigh on agentic use: a large context window, which makes it possible to handle a whole scope without splitting it, and the option of delegating to specialized sub-agents. In practice, the advantage depends above all on your design: structured outputs, control over tools, and observability. A good model that is badly framed produces runs nobody can review.

 

How does Claude compare with ChatGPT?

 

Both answer the same overall need — conversational models and a technical layer to call them — but the gap plays out on the agentic ecosystem. On one side, repository-oriented tools, with Claude Code in the terminal, a large context window and delegation to sub-agents; on the other, an agent mode geared to execution in a browser. The right criterion is therefore what your agent has to touch.

 

How do you create an agent with Claude?

 

Follow a sequence, in this order: define the objective and the indicators (quality, time, cost, success rate); define the authorized tools (read, write, commands, integrations); impose a verifiable output schema; add the guardrails (budgets, approval, rollback); log and test on a pilot scope before any extension. The most frequent mistake is swapping the first two steps.

 

Is Claude Code essential for building a development-oriented agent?

 

No, but it is a decisive accelerator if your agent has to act directly on a repository. The difference is clear: in a chat, you copy and paste the code and you are the execution loop yourself; a Claude Code agent operates in the terminal across the whole project, runs commands, launches the tests and iterates on its own. If your task touches no files, the benefit disappears.

 

How do you structure sub-agents to avoid losing context and making errors?

 

Split by deliverable, not by abstract role: a sub-agent must be judgeable on what it hands back. Between them, impose a short summary always in the same format — done, still to do, decisions, files touched — rather than a transfer of raw context. And favour genuinely parallelizable tasks: each sub-agent restarts from an empty context and consumes its own budget.

 

What guardrails should be in place before authorizing file writes or command execution?

 

Apply a minimal permissions policy: writing restricted to a dedicated branch, execution of routine commands (tests, lint, build) in an isolated environment, since they execute code that comes from the repository, secrets outside the workspace, and a mandatory merge request on protected branches. Add budgets — iterations, duration, cost — and require final evidence: green tests and the list of files modified.

 

How do you assess an agent’s reliability in real conditions?

 

Measure on a set of representative tasks, with a constant protocol. Track at minimum quality (tests, human review, regressions after merge), time (total duration and review time), cost (iterations consumed) and success rate (tasks completed without escalation). The ratio of review time to agent time is the indicator that decides whether to roll out more widely.

 

How do you reduce hallucinations and impose verifiable answers?

 

Do not ask for “correctness”, ask for “provability”. Require internal references — file paths, commands executed, test outputs — and, for any claim that goes beyond the repository, impose an explicit reference or an escalation if the agent can cite nothing. Give it its sources of truth as well: without them, it will compensate with the plausible.

 

Which “French AI agent” use cases are the most realistic in a company?

 

Those whose deliverable is easy to check: technical documentation, test generation, incremental refactoring, classification of incoming requests with structured justification, operational checklists. The value comes from the ability to produce deliverables that French-speaking teams can review, with precise business vocabulary and clear traceability — that is, from your scoping, not from the model’s language.

 

Continue reading

 

  • Your task is not about a repository but about third-party sites and interfaces: a remote browser and taking back control are the subject of the ChatGPT AI agent, not of local execution.
  • You want the method before the tool, or your agent goes beyond the scope of code: the full approach is set out in the guide to create an AI agent.
  • You are stuck on the level of autonomy itself and want the general rule: thresholds, delegation tiers and what is never delegated are covered on autonomous AI agents.
  • Your subject is becoming the contribution chain rather than local execution: branches, reviews and merging belong to the GitHub AI agent.

Discover other items

See all

Next-Gen GEO/SEO starts here

Complete the form so we can contact you.

The new generation of SEO
is on!

Thank you for your request, we will get back to you as soon as possible.

Oops! Something went wrong while submitting the form.