Tech for Retail 2025 Workshop: From SEO to GEO – Gaining Visibility in the Era of Generative Engines

Back to blog

OpenAI AI Agent: What You Assemble, What You Evaluate, What You Authorize

GEO

Discover Incremys

The 360° Next Gen SEO Platform

Request a demo
Last updated on

26/9/2026

Chapter 01

Example H2
Example H3
Example H4
Example H5
Example H6

Two planes not to confuse: the product and the build layer

 

At OpenAI, the word “agent” covers two objects that share neither buyer, nor cost, nor lifespan. On one side a product experience: you describe a task in an interface and watch it run. On the other a build layer — an API, a development kit, callable tools — on which a team assembles a system that will run with nobody in front of the screen. This article deals with the second plane, and it is no side road: more than 2 million developers use the API (Chad Wyatt, 2026), which accounts for 20% of messages in enterprise workflows (Chad Wyatt, 2026). Those readings appear in our ChatGPT statistics.

Two markers save a round trip. What the product runs on its own, without a line of code, belongs to the ChatGPT AI agent mode. And if the ecosystem is not settled, comparing families of tools and setting the requirements of the IT department and the security officer belongs to the choice of an AI agent platform. Here the ecosystem is assumed to be chosen: what remains is what you assemble, what you measure, and on what condition you widen the scope.

 

One-off use or a system that runs continuously

 

The confusion is expensive: an “agent inside ChatGPT” optimizes individual productivity, whereas an agent built on the API has to handle permissions, data, quotas, observability and business rules. Yet both are demonstrated the same way, and that is what misleads: a demo never shows the behaviour on the thirtieth run, on incomplete data, when a source has changed format and nobody is watching.

The deciding question fits on one line: do I need a one-off use, or a system that runs continuously, with evaluations and governance? If the task comes back once a quarter and a human triggers it, the product is enough. If it comes back every week, has to start without being launched and produces an output that other systems consume, you are building — and everything that follows applies.

 

What the confusion costs on the organizational side

 

The cost of the confusion is not technical, it is organizational. Product use asks for nobody: everyone adopts it, drops it, picks it up again. An agent built on the API asks for a named owner, able to answer for what the system did at three in the morning on a Tuesday. It also asks for three skills rarely found in the same person: someone who writes the business rules, someone who holds the integration and the access rights, someone who maintains the evaluation case set.

Maintenance is precisely the item forgotten at scoping. An agent is not delivered once: models change, sources move, business rules shift, and each of those movements can make a behaviour that used to work regress. Budget the replay of evaluations as an on-call duty: periodic, named, independent of whoever built the agent. Otherwise, within a few months the agent turns back into a script nobody dares to modify.

 

What OpenAI supplies for building, and what stays on your side

 

A vendor ecosystem is judged on two lists: what it saves you from assembling, and what it leaves you with anyway. The first is the one in the demos; the second decides your schedule. The maturity of the whole is no longer the question — 1.5 million business customers (Chad Wyatt, 2026) — but a mature ecosystem supplies building blocks, never your system.

 

The blocks supplied: visual building, development kit, tools

 

Four families of blocks stand apart, all resting on the same API, and you need to know which one you are using: they share neither audience nor exit cost. None of them comes with a ChatGPT subscription: product usage and the build layer are billed separately, on separate accounts.

  • A visual build environment: you assemble steps, version them and set guardrails without writing code. Quick to prototype with, and the right medium for having a sequence reviewed by the business.
  • A development kit (SDK) for agents: the same logic in code, calling the same API, with types and tests. That is where the agents you maintain end up.
  • Tools the model can call: sourced web search, search across files, code execution, browser control. These are capabilities, not integrations: they widen what the agent can do, not what it is allowed to do.
  • Connectors to third-party applications, with the associated authentication, access scopes and quotas.

A fifth block joins them, one that goes uncounted because it produces nothing visible: the evaluation layer, which replays a case set and grades execution traces, and which is covered further down. The distinction between the first two families is not cosmetic: a visual sequence reads quickly but exports badly, a coded sequence can be compared and restored. Prototype with the first, ship with the second, and place the switch where the agent gets its first write right.

 

What the ecosystem does not supply, and what stays with you

 

The rest belongs to you, and it is what makes the project.

  • The business rules: what the agent must refuse, what it must escalate, what is never negotiable.
  • The quality of the data consulted: an up-to-date base, with read rights that mirror those of your teams.
  • The evaluation case set: nobody can write it for you, because it encodes what “correct” means at your company.
  • The connection to your systems: recovery on error, queues, choosing what gets logged.

None of those four items is settled by a choice of tool, and none disappears if you change vendor. That is why a prototype convincing in three days takes three months to reach production: the three days consume what the ecosystem supplies, the three months what it leaves.

 

The core of an agent: instructions, output formats, stop criteria

 

An agent in production is not “a longer prompt”. It is an architecture that connects model, tools, data, memory, evaluation, security and supervision. The core rests on three things: stable instructions (rules, objectives, style), output constraints (schemas, formats) and stop criteria — the moment the agent stops deciding and escalates to a human.

Those elements are written before the first line of code, and in black and white: what the agent is allowed to do, what it must ask for, what it must refuse. As long as they are not written, you do not have an agent but a prototype about which nobody can say whether it got things wrong.

Item to write down What it settles Concrete example (B2B) Why it is critical
Objective The single task and its trigger A weekly summary of gaps and the actions proposed Prevents the « jack-of-all-trades » agent
Acceptance criteria What every valid output contains Mandatory citations, prioritized actions Makes the output checkable
Action limits What the agent never does alone No CMS publishing without human approval Reduces the risk of regression
Stop criteria The signal that triggers escalation Missing source, contradictory reference data, out-of-scope request Turns a silent error into a request

 

From the plausible answer to the defensible answer

 

A model always produces an answer; nothing guarantees it is verifiable. Three decisions move the output from plausible to defensible, and each is paid for in constraint:

  • Structured output format (table, schema, plan): more reliable automation — machine reading, tickets, chaining — at the price of an output that is less readable as it stands.
  • Validation rules (thresholds, escalation): fewer errors in production, at the price of a high escalation rate in the first weeks.
  • Defensibility criteria (sources required): quality control that can be made objective, at the price of occasionally empty answers when the source is missing — which is the intended behaviour.

The third is the one removed first when the agent “does not answer enough”. That is the mistake: an empty answer points to a hole in the document base; a plausible answer with no source is a debt you only discover downstream, often at a customer’s.

 

The retrieval strategy decides more than memory does

 

An agent’s performance depends less on “magic memory” than on a retrieval strategy: which sources to query, when, and at what level of freshness. In a company, favour controlled knowledge — internal documents, product reference data, brand rules — and make the Web a complement for current events, comparisons and cross-checking.

Write that rule as a priority and not as a preference: the internal base prevails, the Web serves to complete and, where relevant, to contradict. An agent that queries the Web first will hand you answers that are right about the market and wrong about you, with the same confidence. Building that base — chunking, indexing, freshness — is a project in its own right, and it weighs more on final quality than the choice of model.

 

Evaluate before widening: case sets, graders, trace grading

 

This is where the block-based approach parts company with product use: the measurement setup is part of what you build, on the same footing as the instructions. You do not widen an agent’s autonomy because it “looks good”, but because a replayed case set gives a stable result. The loop has three stages:

  • Define a set of realistic cases (simple tasks first, then edge cases).
  • Measure quality (accuracy, completeness, citations, rule compliance, cost).
  • Optimize (prompts, tools, retrieval strategy), then repeat.

Two instruments tool it. Graders are automatic checkers: to each case you attach what can be observed mechanically in the output — a field present, a format respected, an expected value, the presence of a citation, that of a refusal pattern. They observe, they do not appreciate meaning: an answer can satisfy every criterion and still be wrong, and a grader that delegates judgement to a model inherits that model's errors. Semantic criteria — does the answer really say the right thing — therefore stay arbitrated by a human, on a sample. Trace grading applies afterwards to real runs, put through the same criteria. The first tells you whether the agent holds up on what you had imagined; the second, on what you had not.

 

Writing a usable case set: how many, what variety, who writes it

 

A case set is not a statistical sample: it is a list of situations whose right answer you already know. A few dozen are enough to start, provided the split is right — half nominal cases, half cases that must fail cleanly: missing data, contradictory sources, out-of-scope request, an attempt to obtain an unauthorized action. A set made only of nominal cases validates an agent that has never met reality.

It is written by the business, not by the technical team: the business is what knows that an answer omitting a contractual constraint is wrong, however well phrased. The technical team then turns it into an executable case and a grader. And that set is alive: every incident in production adds a case to it, and that case never leaves. It is the only non-regression you have on a probabilistic system.

 

What the result authorizes, and what sends it back down

 

A score is useless unless it is attached to a decision. So settle, before measuring, what each axis authorizes and what it takes away: without that, you will have a dashboard nobody reads and an autonomy that widens out of weariness. The grid is set once and reread at every widening.

What is measured What it is measured on What the result authorizes What sends it back down a notch
Accuracy Nominal cases with a known expected answer Moving from a reviewed proposal to an automatic draft An error on a case that already passed
Rule compliance Edge cases: expected refusals and escalations Granting the first write right, on a reversible scope An expected refusal that did not happen
Source quality Presence, freshness and relevance of the sources cited Replacing systematic review with sampling An invented source, or an outdated one flagged by nothing
Cost Tool calls and context per run Increasing the frequency or the scope handled Consumption drift with no gain in quality
Latency End-to-end time, nominal cases Moving from manual launch to automatic triggering Going beyond the delay the business accepts

 

The right-hand column is the most important, and it is the one people forget to write: a setup that only knows how to widen is not a control, it is a ratchet. Name who has the right to send the agent back down a notch, without a meeting: that step back must cost less than an incident.

 

The guardrails specific to an agent that browses

 

As soon as an agent reads the open Web and accesses your data at the same time, the risk surface changes in kind: the danger is no longer only that it gets things wrong, it is that it obeys someone other than you, with nothing in its output to flag it.

 

Prompt injection: what it is, and why a better prompt is not enough

 

A prompt injection is a malicious instruction hidden in content the agent consults: a page, a document, a message received. The agent does not structurally distinguish what it is asked from what it reads; a well-placed text can therefore make it disclose content, bypass a rule or trigger an action. That risk is specific to agents that look for information outside, and it is not fixed by writing firmer instructions.

Three countermeasures hold: explicit confirmations before any consequential action, switching off connectors not needed by the current task, and minimal permissions on the ones that remain. The logic is constant: you do not try to stop the agent reading hostile content, you make obeying that content useless. Restricting write rights reduces the risk surface, it does not remove it: an agent that only reads remains exposed to injection, to the disclosure of internal content in its answer and to the exfiltration of data slipped into a tool query or into a link it follows. Writing adds to that the unintended action, carried out in your name.

 

Minimal permissions, human approval, traceability

 

Three guardrails are set at scoping, not after the first incident:

  • Minimal permissions: read access by default, writing only on low-risk scopes.
  • Human approval: mandatory for emails, purchases, irreversible changes and sensitive content.
  • Traceability: log the tools called, the sources consulted, the decisions and the results.

Traceability is the only one of the three that cannot be caught up later: a permission can be narrowed at any time, an approval can be added, but a trace not written at the moment of execution will never exist. Decide what gets logged before go-live, and check on a real case that you can answer the question that follows every incident: which source did the agent consult, and why that action?

 

Going to production: connecting, supervising, holding costs

 

A good deployment favours repeatability and supervision: you start small, you measure, then you widen. Three connection rules cover almost every case:

  • Expose your data for reading with minimal permissions.
  • Break actions down into atomic operations: create a ticket, export a report, propose a draft.
  • Add webhooks and approvals as soon as the agent leaves “advice” and enters “execution”.

The third rule carries the only risk threshold worth memorizing: as long as the agent advises, the worst outcome is lost time; the moment it executes, it is a system state to undo. Breaking actions into atomic operations makes each one separately reversible, and that is what makes the first write right acceptable.

Three rules then frame the scale-up:

  • Supervision: dashboards and alerts on consumption drift, weak sources and repeated errors.
  • Breakdown: avoid “monolithic” agents that do everything in one go and cannot be diagnosed.
  • Autonomy: increase it only after stable evaluations, never after a good week.

That leaves cost, and it is not ChatGPT's: a subscription to the product covers neither the API, nor the development kit, nor the tools called, nor the evaluations. Pricing for that layer has to be thought of as a mix: volume of requests, size of outputs, tool calls — web, files, code — and human supervision. It is the item most often absent from estimates, and the only one that does not fall on its own. Amounts and tiers move too fast to be copied out here: read them off the publisher's official pricing page, for the build layer and for the product separately. Usage is moreover bounded by rate limits and quotas: check what happens when a quota runs out mid-run, because the answer tells you whether the agent stops cleanly or leaves a job half done. Finally, size it knowing that what clears the deployment stage settles in: the twelve-month retention rate of the enterprise offering reaches 88% (First Page Sage, 2026). Your agent will still be there in a year, with its debt.

 

FAQ on OpenAI AI agents

 

How do you use the OpenAI API to create an agent?

 

You combine a model, instructions and callable tools, all orchestrated by a sequence you version. The robust approach has four stages: scope the use case and the action limits, connect the sources (internal files, Web, connectors), define constrained output formats, then put the evaluations in place before increasing autonomy. The order matters: an agent evaluated after the fact cannot be corrected, it has to be rebuilt.

 

What is the OpenAI agent platform?

 

It is the set of building blocks offered around the API: a visual environment to assemble and version a sequence, a development kit to do the same thing in code, built-in tools (sourced web search, files, code execution, browser control) and connectors to third-party applications. On top comes an evaluation and optimization layer: custom graders and grading of execution traces.

 

What is the pricing?

 

There are two of them, and confusing them throws off any budget. ChatGPT's is a per-user subscription, with included volumes and an overage mechanism; it opens no rights on the build layer. That of the build layer — API, development kit, tools called, evaluations — follows usage: number of requests, length of outputs, tool calls. The amounts are not to be quoted from memory: they are to be checked on the official pricing page, which itself distinguishes the two. To those items add the cost of human supervision, which stays the most stable and the most often forgotten in initial estimates.

 

What are the capabilities of OpenAI agents?

 

An agent built on this layer can reason and then act through tools: search the Web with sourced answers, read files, execute code, control a browser, and call business applications through a connector. It keeps a task context between steps, which lets it chain without starting over. Those capabilities describe what it can do; what it is allowed to do depends entirely on the permissions you grant.

 

What is the difference between an OpenAI agent, a conversational assistant and a tooled workflow?

 

A conversational assistant answers requests, but guarantees neither execution nor traceability. A tooled workflow chains predefined steps, but stays rigid as soon as the case leaves the expected path. An agent chooses and orchestrates tools according to context, iterates, and asks for approval before important actions. The practical difference is not intelligence: it is the right to act, and what gets logged about it.

 

When should you move from a single agent to multi-agent orchestration?

 

When a single run has to carry incompatible responsibilities: searching and citing, executing, checking quality, verifying compliance. The signal is simple: if you are adding guardrails and approvals to the point of making the agent “heavy”, split it into sub-agents with structured inputs and outputs, and therefore separately evaluable. Until that point, multi-agent adds coordination without adding quality.

 

What minimum guardrails limit errors, risky actions and data exfiltration?

 

Three are enough to start, and they are not negotiable: an explicit approval before any irreversible action (purchase, email, publishing), minimal permissions with only the strictly necessary connectors, and a retained trace of the tools called, the sources consulted, the decisions and the outputs. The risk of prompt injection makes the second point decisive: the fewer rights the agent has, the less a hostile instruction can exploit.

 

How do you evaluate an agent before deploying at scale?

 

Build a case set from real data, edge cases included, then attach to each one what a grader has to observe in the output. Measure five axes: accuracy, rule compliance, source quality, cost, latency. Replay after every change, and complete it with grading of real execution traces. Only widen the rights after several stable replays, never after a single good result.

 

Continue reading

 

  • You want the method before the tool: scoping, writing the rules, testing, deploying and governing is set out in the guide to create an AI agent.
  • Your sticking point is the document base rather than the model: chunking, index and freshness are covered on the RAG AI agent.
  • Your single agent has become too heavy to keep under control: coordination patterns, conflict arbitration and replay belong to AI agent orchestration.
  • The blocks are chosen and what remains is connecting them to your systems: connectors, webhooks and recovery on error are the subject of AI agent integration.

Discover other items

See all

Next-Gen GEO/SEO starts here

Complete the form so we can contact you.

The new generation of SEO
is on!

Thank you for your request, we will get back to you as soon as possible.

Oops! Something went wrong while submitting the form.