Tech for Retail 2025 Workshop: From SEO to GEO – Gaining Visibility in the Era of Generative Engines

Back to blog

Using an AI Agent in VS Code

GEO

Discover Incremys

The 360° Next Gen SEO Platform

Request a demo
Last updated on

26/9/2026

Chapter 01

Example H2
Example H3
Example H4
Example H5
Example H6

The extension is in place, and the question is no longer how to install it: it is what happens when you switch to “agent”. Using an AI agent in VS Code does not turn on the power of the model but on the frame you set before clicking: a written scope, acceptance criteria, a list of what it never touches, and a readable diff. For the general method — framing, level of autonomy, specification, acceptance testing — knowing how to create an AI agent gives the frame within which everything below applies.

 

Using an AI agent in VS Code: agent, chat or completion

 

Three things live side by side in the same editor and all get called “the AI”. Completion writes on from your keystrokes. Chat explains, proposes, answers. Agent mode takes a high-level objective and runs it in several steps. The difference is not one of power but of surface: the first touches a line, the last touches a repository. So the right question is not “can the AI do it?”, but “can I check the result quickly?”.

 

What “agent” means in the VS Code agents ecosystem: autonomy, tools, sessions, multi-file

 

An agent in a development environment is not a “chat that codes”: four words separate it from the rest. Autonomy: it chains actions without asking your permission at every step. Tools: it does not only produce text, it reads files, runs tests, calls a terminal. Sessions: its work has a start, an end and a state you can inspect. Multi-file: it acts on a coherent set, not on the open file.

Where an assistant offers a suggestion, the agent runs a loop: it acts, observes the result — tests, static analysis, logs — then starts again. Those four properties are not, however, found in every extension or in every mode: some offer only completion, others a chat with no write right, others a run with no planning and no inspectable session history; the same word covers different capabilities from one tool and one version to the next. So check extension by extension what actually exists, instead of deducing it from the word “agent”. Where a single list of sessions is offered — runs on your machine, on the command line or remotely — that is the view to look at before judging what happened. The table below sums up what each level does, and what it is checked against, where it is available.

Level What it does What it is not trusted with How you check
Completion Completes as you type, on repetitive patterns Business logic you have not yet written yourself Immediate review of the accepted line
Chat Explains an error, offers a lead, compares two options Any direct change: it proposes, you apply Checking against the documentation and the real code
Plan mode Splits the objective into steps, lists the files affected and the risks Execution itself, for as long as the plan has not been reviewed Reading the plan: scope, order, stopping points
Agent mode Changes several files, runs commands, iterates on failures Secrets, production configuration, writing outside the repository Full diff, tests run, session log

 

When a simple assistant is enough, and when the agent becomes useful

 

An assistant is enough when you already know what to change and want to go faster on a short unit of work: a function, a query, a test. An agent becomes useful as soon as the task crosses the project: long tasks, refactorings spread over several modules, migration of an interface, writing then running a battery of tests. The tipping criterion is not the intellectual difficulty of the task, it is the number of places that have to be touched before it is finished.

In the development cycle, the agent sits mainly between the specification and the code: it turns an intention into a plan, then into verifiable changes. It also adds value downstream, on review and on runbooks: diagnostic procedures drawn from the logs, documentation updates, deployment checklists. The agent speeds up execution, product responsibility and quality stay human.

 

Running an agent session

 

An agent performs well when you feed it like a contributor arriving on the project: an objective, constraints, and a clear definition of “finished”. A badly opened session cannot be rescued along the way: it produces a large diff you will have to review in full, or accept out of fatigue. Before any change, require an output format and a sequence: plan, files affected, commands run, summary of the changes.

 

What gets written before the agent is allowed to change anything

 

Four elements are set at the start of the session, and they are carried over from one session to the next. The project context — conventions, folder structure, versions, logging style — stops the agent inventing the house rules. The instructions say what it has to produce, the limits where it stops, the acceptance criteria and the definition of done how you will know it is finished.

  • Objective: the result expected on the user side, not the way to get there.
  • Constraints: technical frame, versions, style, performance and security requirements.
  • Acceptance criteria: tests that have to pass, expected behaviour, edge cases covered.
  • Sequence: a plan before writing, then implementation in steps, each one verifiable.

The test is simple: a colleague who does not know the task should be able to say, on reading the acceptance criteria, whether an output is compliant or not. If the answer depends on your judgement, the agent has no way of stopping in the right place.

 

Setting autonomy by the risk and the complexity of the task

 

The setting is redone for every task, and it depends on two variables. The surface of change first: the more files are touched, the more you have to reduce autonomy and multiply validation points. Risk next: an area covered by solid tests tolerates a wide run, an area with no tests tolerates none. A highly autonomous agent on untested code does not go faster; it moves the verification work to a moment when it costs more.

Four guardrails are set at the level of the session itself. Read-only first: the first pass produces an analysis and a plan, never a write. Dedicated branch: the agent works alongside, with atomic, reversible commits. A branch isolates the history, nothing more — it restricts neither access to the files on the machine, nor the commands run, nor what goes out in the context. Systematic review of the diff: no unexplained change is accepted. Stop conditions, finally: if a test fails twice in a row, the agent stops and asks for clarification instead of attempting a third fix.

 

What an agent must not touch in your repository

 

The list of prohibitions is written before the first session, not at the moment the agent asks for permission: at that instant, you want the task to move forward, and you approve. That is the mechanism by which a scope is lost. This list is short, it lives in the repository’s instruction file, and it is reviewed as a team. It has two halves: what the agent may not change, and what it may not read. One clarification, before going through it: an instruction file is an instruction addressed to the model, not a technical barrier. It stops neither the reading of a secrets file present on the disk, nor the running of a command outside the repository; it is text, which the model can neglect and which content read elsewhere can contradict. What does stop things belongs to the system: rights of the account running the agent, secrets taken off the machine and placed in a vault, mandatory approval of terminal commands, an isolated execution environment. Instructions reduce how often deviations happen, barriers reduce their reach: you need both.

 

Files, secrets, dependencies and commands

 

The file scope is declared positively: the folders the agent may write in, and nothing else. An exclusion list is forgotten as soon as a new file appears, an inclusion list is not — provided a mechanism enforces it: declared in a mere instruction file, it remains an intention that nothing compels anyone to respect. Excluded by default are the environment files and anything holding a secret — tokens, credentials, keys — staging and production environment configurations, migration files already applied, and generated artefacts that a tool rebuilds.

  • Dependencies: the agent proposes an addition, it never installs one on its own. A dependency enters the project through a human decision.
  • Running commands: allowed on verification commands — tests, formatting, static analysis — subject to explicit approval for everything else, and never on a command that reaches a system outside the workstation.
  • Writing outside the repository: forbidden, without exception, and the ban is set in the execution environment rather than in an instruction, failing which it is only a wish. An agent that changes your global configuration produces a change the diff will never show.
  • History and protected branches: no history rewriting, no action on the main branch. The work comes out as a proposal and is reviewed before it goes in.

These four rules protect what the diff does not show: a dependency installed silently, a command outside the scope or a file written next to the repository all escape review.

 

What does not go into the context

 

The second half of the list concerns reading, and it is the easiest to forget: it leaves no trace in the repository. One point needs stating before all the rest: an agent that runs in your editor is not for that reason a local agent. Unless a model is installed on your machine, every step transmits to a remote service what it has read — code extracts, paths, command outputs, error messages, and sometimes the content of files you have never opened. The reflex of opening the context wide so “the agent understands better” is what sends data off your machine. The risk is not theoretical: uncontrolled entry of sensitive data concerns 4.7% of enterprise users (Chad Wyatt, 2026) — an order of magnitude, which varies between organizations and tools, but enough to justify a written rule. These benchmarks and their variants are gathered in our review of ChatGPT and generative AI statistics.

Give the minimum sufficient context: conventions, objective, paths of the relevant files, the extracts needed — not the whole repository on the grounds that it is indexed. Left outside are secrets and environment files, personal data and customer datasets, production log exports and the contractual documents lying around the project. When a test case needs realistic data, synthetic data replaces it. If your internal policy requires it, running the model on your own machine becomes the only acceptable answer, and that is decided before the project.

 

Assessing an extension before authorizing it

 

An AI extension often has access to the workspace, sometimes to the terminal, and can trigger actions that go well beyond “a prompt”. So the assessment happens beforehand, on five points, and it ends in a team decision: authorized, authorized under conditions, or refused. An extension “tried out to see” on a real repository has already been authorized, without anyone having decided so.

 

Five points to check before installing an extension

 

The data sent and the permissions say what the extension can see and do. Logging says whether you will be able to reconstruct what happened. Reversibility says whether you will be able to go back, and the supply chain says who maintains the code you are installing and how often. Those last two are the most discriminating, and they are the ones that get skipped.

Checkpoint Question to ask Expected decision What disqualifies
Data sent Full code, extracts, metadata, logs? Sharing rules per repository and per folder No clear answer on what leaves the machine
Permissions Read-only, write, terminal, network? The minimum level needed, granted explicitly Permissions that cannot be separated from one another
Logging Can the actions and decisions be audited? A session log usable after the fact Actions invisible once the session is closed
Reversibility Can a batch of changes be cleanly undone? A rollback possible without rebuilding by hand Changes applied with no restore point
Supply chain Who maintains the extension, and how often? A policy of authorized extensions, reviewed periodically Unknown maintainer, opaque updates

 

The robustness limits no extension fixes

 

Four limits come from the nature of the models and are found in every extension. Error handling: an agent that fails tends to try a variant rather than question its starting assumption. Multi-file consistency: a change that is right in each file taken on its own can be wrong at module level. Context limits: beyond a certain size, the agent loses elements it had in fact read, and nothing tells you. Non-deterministic behaviour, finally: the same request, twice over, does not produce the same diff.

Models remain probabilistic: they can produce a convincing but incorrect result, or invent an interface that does not exist. Indexing the workspace reduces how often the problem occurs without removing it, and makes answers more plausible, and therefore harder to challenge on a quick read. So for any multi-file change, require an explicit list of impacts and an automated run of the checks, before looking at the code.

 

Cutting down the iterations

 

Time lost with an agent is rarely lost all at once: it is lost in back-and-forth. A good agent prompt is a small technical brief. And it explicitly asks for a verifiable result — commands to run, tests to add, success criteria — rather than a satisfying one.

  • Objective: “add an endpoint X with input validation and explicit errors”, not “improve the API”.
  • Constraints: versions, conventions, folder structure, logging style.
  • Expected formats: list of files touched, diff, tests, migration note.
  • Minimal examples: one input case and its expected output, as short as possible.
  • Anti-examples: “redo the whole API” with no scope and no criterion — and, more useful still, an extract of what you do not want to see in the result.

The anti-example is the most profitable of the five, and the least used: it costs two lines and removes the class of errors you have already hit twice. The second profitable reflex: have an action plan produced before any writing — steps, files, risks — then move forward in batches. Splitting into tasks has a mechanical effect on review: four diffs of thirty lines get read, one diff of a hundred and twenty gets skimmed.

Between the batches, set validation checkpoints: the agent stops, shows what it has changed and how it checked it, and waits. The full sequence fits in four stages — implementation plan, first step and a small diff, tests then fixes, final recap stating what changed, how to check it and how to roll back. That last point lets you abandon a failed session instead of repairing it.

 

Demanding evidence: tests, diff and artefacts

 

An agent produces plausible code. Your job is not to judge that plausibility by reading, but to demand evidence that does not depend on your appreciation. Three families are enough: tests that fail when the code is wrong, a diff that can be read line by line, and artefacts that explain why a decision was taken. The first two protect today’s version, the third the one somebody will pick up in six months.

 

Having the scenarios, the edge cases and the non-regression criteria produced

 

To debug, first have the assumptions made explicit: probable cause, files involved, way to reproduce. An agent that announces a fix without having reproduced the problem is fixing something else. Then impose a three-stage test strategy — nominal case, edge cases, non-regression — and ask for the edge cases before the fix: produced afterwards, they are written to pass.

The non-regression criterion is the one that gets forgotten and the most discriminating: it checks that what worked before still works, which neither the diff nor the review shows. Finally, demand verifiable fixes: the command run, the result observed, and the explicit link between the cause identified and the change made.

 

Reviewing the diff, and what the agent must leave behind

 

Your best quality assurance remains a readable diff and atomic commits. An agent that produces an unreadable diff is not reviewed: it is approved. Ask for conventions respected, commit messages that explain the intention rather than the mechanics, and a mention of the risks when the change touches a sensitive area. Traceability and auditability are not added afterwards: they are what will let you say, in three months, why this line exists.

Then come the artefacts, which the agent maintains well and which nobody updates spontaneously: README, architecture decision records, changelog, runbooks. They are reviewed like code. Demand the same of them as of the rest: runnable examples, extracts tied to the real code, stable conventions and the known pitfalls. If the subject becomes writing the loop that does all this yourself — project structure, explicit state, deterministic tools, tests — then it is a Python AI agent you are building, and that is no longer the same job: here, an agent writes code for you.

 

Knowing whether it actually saves you time

 

The gain is greatest on standardized tasks: generating skeletons, writing tests, documentation, mechanical refactorings. There the expected result is known in advance, verification is quick, and your added value was nil. The hidden cost shows up elsewhere, when the output is hard to maintain: style inconsistencies, needless abstractions, quiet duplication, technical debt nobody decided to take on. The tipping point is almost always the same — review time exceeds the writing time saved.

This judgement is not made on impression. Four indicators are enough to settle it, with no particular tooling.

  • Review time: the review time for a change produced by the agent, compared with an equivalent change written by hand. If it is higher, the writing gain has already been cancelled out.
  • Share of rejected fixes: the proportion of proposals abandoned during a session. Beyond a third, it is the framing of the sessions that has to be reworked, not the tool.
  • Edits after merge: the corrections made in the following days, on code the agent had produced. It is the most honest indicator, because it counts what review let through.
  • Observed debt: the areas of the project the team avoids touching. If they appear where the agent wrote a lot, the speed gain was borrowed, not earned.

A team that collects those four points over a month gets a documented decision on the tasks it hands to an agent and those it keeps. And it will not be the same on a new, well-tested project as on an old system whose dependencies nobody fully knows any more.

 

FAQ on AI agents in VS Code

 

How do you code faster with AI in VS Code without degrading quality?

 

Go fast on what can be checked fast: code skeletons, tests, documentation and mechanical refactorings. Impose a short loop — plan, small diff, tests, review — rather than one big change in a single pass. And use the agent as a reviewer and a generator of edge cases as much as an author: that is the role in which it reduces technical debt instead of increasing it.

 

How do you create an agent in VS Code (and which limits should be set)?

 

You can use the built-in modes — ask, plan, execute — and define custom agents by giving them a role, a list of available tools and a model. Set the limits before the run: authorized file scope, explicit prohibitions (secrets, production configurations), mandatory tests. The riskier the change, the more explicit approvals you require during the session.

 

How do you use Copilot in VS Code, including Copilot agent mode?

 

Open the assistant’s conversation view, switch agent mode on in the settings if your organization has not already done so, then move to that mode for multi-step tasks: multi-file changes, commands, tests. Always start with a request for a plan, then let implementation proceed in steps with validation points between each.

 

Which AI extensions support agents (and how do you assess them)?

 

Assess an extension on five axes: data sent, permissions, logging, robustness — errors and multi-file consistency — and reversibility. Favour those that make their actions auditable and that fit cleanly with your tests and your conventions. If you have to choose between two, pick the ability to prove — tests, logs, diff — over the ability to produce quickly.

 

What is the difference between an agent, a chat and autocompletion in VS Code?

 

Autocompletion speeds up typing on known patterns. Chat helps you understand and decide, but it changes nothing: you apply it yourself. The agent carries out a complete task in several steps, with tools, and touches several files. The right choice depends on the surface of change and on your ability to check the result quickly.

 

Which guardrails should be switched on before letting an agent change several files?

 

Work on a dedicated branch, require atomic commits and make tests mandatory before any merge. Declare the file scope positively, exclude secrets and sensitive configurations, and forbid writing outside the repository. Add stop conditions: on a test’s second failure, the agent stops and asks for clarification.

 

How do you give an agent the right project context without exposing sensitive data?

 

Give the minimum sufficient context: conventions, objective, paths of the relevant files, the extracts needed. Exclude secrets, personal data and production log exports, and replace real data with synthetic sets in the test cases. If your internal policy requires it, running the model on your own machine becomes the only acceptable answer.

 

How do you avoid plausible errors and improve verifiability (tests, logs, evidence)?

 

Ask for evidence at every step: the command run, the result observed, the explicit link between the cause identified and the fix applied. Have the edge cases produced before the fix, never after, and systematically add non-regression tests. An output with no evidence is treated as a proposal, not as a result.

 

Can an agent be run locally, and when is that preferable?

 

Distinguish two things. An agent can run on your machine, including in the background from a command line, while calling a remote model to which it sends its context: that is the most common case, and it keeps nothing at your end. Local in the strong sense additionally assumes a model installed on your machine. That is the one that becomes necessary when the repository holds sensitive code, or when your compliance rules limit sending context to an external provider. It is paid for in operating load: dependencies, performance, reproducibility of results.

 

Can an AI agent in VS Code help maintain AI features in production?

 

Yes, above all to speed up putting repetitive patterns in place — wrappers, validations, tests, documentation — and to keep multi-file consistency. But the agent does not replace architecture: demand traced decisions, tests and an operations runbook, failing which you will be blindly maintaining a system nobody designed.

 

Continue reading

 

  • The diff is reviewed and approved on your machine, and what comes next happens on the remote repository: automated review, triggering and the continuous integration chain belong to the GitHub AI agent.
  • You are no longer deciding for yourself alone but for a team, and the question becomes one of licences, account rights and policies: that is the ground of the Copilot AI agent.
  • You want to run the agent somewhere other than in your editor: models, vendors and automation tools are compared on an AI agent platform.
  • The code must not leave your machine, and that constraint comes before convenience: hardware, models and operating load are compared on the local AI agent.

Discover other items

See all

Next-Gen GEO/SEO starts here

Complete the form so we can contact you.

The new generation of SEO
is on!

Thank you for your request, we will get back to you as soon as possible.

Oops! Something went wrong while submitting the form.