26/9/2026
You have a Python file that calls a model, works one time in three, and that nobody but you can read. The question is no longer what an agent is: it is what shape your code has to take so that it can be read, tested, shipped and audited. Writing a Python AI agent is not decided by the finesse of a prompt, but by four ordinary things: a tidy project, an explicit state, deterministic tools and tests that stop a regression. For what is decided upstream — framing, level of autonomy, specification, decision protocol — knowing how to create an AI agent gives the frame that the code below implements.
From script to agent: what changes in your code
An AI agent is not an improved “chat”. It is a system designed to take decisions and carry out autonomous actions in a given environment, often through external tools — database, web search, programming interface — where a conversational agent stays confined to generating text with no real action. The difference is not the quality of the text produced: it lies in acting, and therefore in the obligation to prove what was done.
Your practical criterion is simple: if your code does anything other than return an answer — create a report, run a query, produce a deliverable, trigger an action — and it evaluates itself, you are already in agent territory. A script runs a sequence of instructions. An agent maintains a state, chooses its actions according to an objective, relies on external tools, then checks its results before deciding what comes next.
Before writing a single line, set down a shared vocabulary. In production, confusion between model, tool, agent and orchestration is paid for in bugs, in drift and in invoices.
- The language model: it generates, classifies, rephrases. It is the only building block you accept as probabilistic.
- The tool: it carries out a deterministic action, verifiable and replayable identically.
- The agent: it plans, chooses the tools, checks the result obtained and decides what comes next.
- Orchestration: it paces, triggers and supervises the runs.
This split carries the whole argument of this article: plug a language model in for the probabilistic part — planning, rephrasing — with everything else staying verifiable: data, calculations, exports. If a rule always has to give the same result, it is written in Python, not in a prompt.
Three shapes come back almost every time, with the same building blocks: the research and synthesis agent, which collects and delivers a structured summary with its sources; the analysis agent, which turns raw exports into prioritized decisions without guessing — it computes, then explains, and keeps the traceability; the controlled production agent, which chains brief, writing, quality checks and delivery, its core being the pipeline of checks and not the model.
Structuring the project before writing the loop
An agent you cannot test is not shipped: it is demonstrated, which is not the same thing. Testability is not added afterwards, it is decided at the moment you create the first folders. The rule that governs the rest fits in one line: nothing that changes from one environment to another — threshold, identifier, path, quota — belongs in the code that decides.
Four folders, and what each one holds
The minimum structure fits in four directories. This is not a matter of style: each one corresponds to a category of change, and therefore to a distinct reason to modify the project. A file that could fall into two of them signals a badly split responsibility.
- Configuration: environments, thresholds, business rules, mappings. What changes without shipping code again.
- Connectors: the clients for external interfaces, with each provider’s quotas, timeouts and errors.
- Agent: the planner, the executor, the validators and the action policy. It is the only folder where decisions are taken.
- Tests: fixed cases, non-regressions and simulated errors from external providers.
Separate the business code from the connectors clearly: it is that boundary which lets you test the agent’s logic with no network, by substituting a connector with a frozen dataset. An agent whose loop cannot be run offline has no tests: it has real calls that somebody reruns by hand.
What does not live in the repository: secrets, configuration, test sets
Move secrets out into a dedicated manager, never into the repository, and separate them by environment. A hard-coded identifier survives reviews — it looks like an ordinary variable — and ends up in the version control history, where it is never really erased. Configuration follows the same logic: loaded at start-up, with an immediate failure when a mandatory value is missing.
Test sets, for their part, do live in the repository: that is why they have to be synthetic. A real export embedded “for testing” brings customer data into a project cloned onto several machines. If your editor itself hosts an agent that reads that repository, the same rules apply to it, with the added question of what goes into its context: what a VS Code AI agent is allowed to touch comes up at the same moment, and is not to be confused with the agent you are writing here.
Writing the agentic loop
A solid agent looks less like a demo than like a controlled loop: explicit state, deterministic actions, evidence, then a decision. The objective is twofold — make errors visible early, and make actions auditable afterwards. The classic trap has a name: the model, tool, model chain that never stops, because no step can say whether it has made progress. You get out of it through a structured state and explicit transitions, not through one more instruction in the prompt.
The four stages: plan, action, observation, decision
Each iteration goes through four stages, and each one produces an inspectable object. That is what separates a loop from a series of calls: at the end of every turn, you know where the agent stands without rereading its outputs.
- Plan: select one single candidate priority action and produce a strict structured output, with its assumptions and its success criteria.
- Action: execute through a connector, with a unique operation identifier, so that a replay does not produce a duplicate.
- Observation: collect metrics and state, before and after, then check assertions rather than ask the model whether everything is fine.
- Decision: continue, correct, escalate to a human, or stop because a threshold has been reached.
The pattern that makes this workable: model the state in a serializable object, then write a utility that replays a run from the logs. The exchange format between the planner, the connectors and the validators is handled as a versioned contract: any input or output that does not conform to the schema is rejected before going any further.
What stops the loop from spinning
A Python agent becomes dangerous when it acts without limits. Frame its autonomy like a budget, with ceilings written in the configuration: number of actions, cost, duration, and stop rules that apply without arbitration.
- Human supervision: mandatory validation before any write — publishing, deletion, sending, updating a third-party system.
- Stop criteria: no evidence, no action; insufficient confidence, escalate; three iterations with no measured progress, stop.
- Action budget: cap network calls, planning loops and the size of the context passed on.
These three guardrails are checked in code, not in an intention. A stop criterion that depends on the model’s own judgement of its confidence is not one: write it as a condition on an observable value — evidence collected, gap between two sources, mandatory field missing — and the stop becomes reproducible.
Tools: making deterministic whatever can be
Your tools have to be more reliable than your model. In other words: everything touching the truth — data, calculations, extraction, formatting — goes through deterministic functions, never through free generation. It is the only way to get the same result twice on the same input. A tool is designed as a narrow function: one responsibility, typed parameters, an output validated by a schema, an explicit error when something is missing.
The Python ecosystem covers the essentials without asking the model for anything. Data access goes through the standard libraries or your engine’s drivers, analysis through the tabular and numerical computing libraries, output through a charting library and a template engine that separates data from presentation. That last point is not cosmetic: a report whose formatting is generated at every run cannot be compared from one version to the next, and therefore cannot be tested.
Before authorizing an action, set three guardrails in the tool itself. Read-only by default: any function that writes is declared as such and requires explicit authorization. Input validation: an out-of-range parameter makes the call fail, it does not make it deviate. Idempotency: the same operation replayed leaves the system in the same state. When the agent has to answer or decide from internal documents, the tool is no longer a simple access path: chunking the corpus, the retrieval strategy and the citing of sources then belong to the RAG AI agent, and that building block is designed in its own right.
Logging, tracing, versioning
In a company, an agent with no observability is unusable, whatever the quality of its outputs. You have to be able to answer three questions, months after the run: what did it do, with which data, and on which rule did it decide? None of those answers can be reconstructed from memory. They are prepared at the moment you write the loop, and they cost a few lines around each tool call.
- Structured logs: action, input, output, duration, error, estimated cost — in a machine-readable format.
- Traces: the call chain and a unique correlation identifier per run, propagated all the way into the connectors.
- Versioned prompts: fingerprint, date, author, change log. A prompt is code, it is reviewed and compared.
- Decision log: stops, escalations, refused actions, with the rule that triggered them.
The fourth line is the one that gets forgotten, and the only one an auditor cares about. Logging actions shows what happened; logging decisions shows why. Without versioning of prompts and rules, you will not be able to say what changed between a correct run last week and a failed one today — and you will spend that time changing the prompt at random.
One last criterion says whether your observability is enough: you have to be able to replay a complete run from the logs alone, without reopening the original systems.
Framework or in-house implementation
An assembly library has to speed things up without hiding the decisions. What costs you most is neither writing time nor the licence: it is unreadability. A “magic” architecture, which stacks abstractions to the point where you lose sight of the plan, action, evidence, decision path, becomes impossible to debug, to test and to secure. The right tool is the one you can run, test and audit at scale.
The criteria that decide, and their warning signals
Five criteria are enough to settle it, and each is checked on a real case rather than on a documentation example. The table gives, for each, the signal that should make you step back and its consequence in the code.
When an assembly library helps, and when it costs
The decision rests on the shape of the problem, never on a tool’s popularity. An assembly library is worth using if you are doing document retrieval, tool routing or modular pipelines. It is to be avoided if your need fits in three tools and two rules: a simple implementation will be more robust and easier to prove.
A second case justifies it: when your problem breaks down into roles — research, analysis, writing, quality control — a role-oriented structure avoids the monolithic agent that does everything and becomes impossible to make reliable. The right pattern then separates an analysis agent, which handles data and calculations, from an editing agent, with a validator that blocks any claim without evidence.
Then comes the most frequent case, wrongly set aside. If you have a clear scope — generating a report from exports, for instance — writing it in house is often the best choice: you gain on readability, on control and on the ability to prove what the agent does. The cost of a dependency shows up at the first version bump that breaks an abstraction you depended on without knowing it.
Acceptance testing before shipping
Test it like a product, not like a notebook. The question is not “does it work” — a demo only shows the nominal path — but “what happens when it does not”. Going into real conditions turns on three subjects: reproducibility, non-regression and cost. That is the part the team receiving the agent cares about: they are not judging your loop, but they know how to read an acceptance criterion and refuse a delivery that does not meet it.
Case sets, acceptance criteria and non-regression
Acceptance testing has three parts, gathered in a document reviewed at every change of prompt, rule or connector. The case sets: normal cases, sensitive cases, missing data, quotas reached. The acceptance criteria: rate of valid outputs, escalation rate, rate of actions correctly refused. Non-regression: same input, same decision, within an acceptable range. The last is the most discriminating and the first to be forgotten: an agent whose decision varies on an identical input is not ready.
Packaging, environments and reproducibility
Reproducibility is a requirement, not a convenience: without it, no test proves anything, since nothing guarantees the run environment will be the same tomorrow. Three moves are enough. Isolate execution in a virtual environment dedicated to the project. Freeze the dependencies in a lock file, indirect dependencies included. Move the configuration out, so that the same artefact runs in development, in staging and in production without a line of code changing.
Add a minimal install check: an entry point that imports the key libraries, checks read-only access to the connectors and exits with an explicit error if a mandatory variable is missing. It costs ten lines and turns a half-day diagnosis into a readable message.
Latency, context and cost per run
Your cost is not only financial: it is also latency, which decides what your users will agree to wait for. The discipline comes down to three levers: reduce the context passed on, limit iterations, and prefer a deterministic tool to one more question put to the model. The last is the most profitable and the least applied: a check written in code gives a stable answer, where rephrasing costs one more call and guarantees nothing.
So set budgets — number of calls, context size, duration — cache the retrievals, and replace re-generation with deterministic checks. Measure the cost per run and cut the loops beyond a defined threshold, set in the configuration and logged when it fires: that is the only way to know whether an agent has become more expensive because it handles more, or because it is going round in circles.
FAQ on Python AI agents
How do you create an AI agent with Python?
Start by creating a controlled loop: explicit, serializable state, a short plan, execution through deterministic tools, then verification and decision. Lay the project out in four folders — configuration, connectors, agent, tests — before writing the first function. Only then plug a language model in for the probabilistic part, planning and rephrasing, with everything else staying verifiable: data, calculations, exports.
Which frameworks should be used?
Choose according to the shape of your problem, not according to a tool’s reputation. An assembly library is justified for document retrieval, tool routing or modular pipelines; a role-oriented structure is justified when the work splits into research, analysis, writing and control. If your need fits in three tools and two rules, an in-house implementation will be more robust and easier to audit.
Which are the best tools?
The best tools are the ones that make the agent deterministic where it has to be: data access through the standard libraries or your engine’s drivers, analysis through the tabular and numerical computing libraries, output through a template engine separating data from presentation. The selection criterion is not feature richness, but the ability to replay the same call and obtain the same result.
What is the difference between a Python agent and a simple script?
A script runs a sequence of instructions. A Python agent maintains a state, chooses its actions according to an objective, relies on external tools, and checks its results before deciding what comes next: continue, correct, escalate or stop. The tipping criterion is practical: if your code produces a deliverable or triggers an action, and it evaluates itself, you are already in agent territory.
How do you avoid hallucinations and enforce verifiable outputs?
Impose a single rule: no claim without evidence. Use tools to extract and compute rather than asking the model, and force an output format containing citations, dates and extracts. Validate every output against a schema before passing it to the next step. With no source available, the agent has to ask for the missing data rather than fill in: that is a performance choice, not excessive caution.
When should persistent memory or RAG be added?
Add persistent memory when the agent has to follow a case across several runs: history, decisions taken, exceptions granted. Add document retrieval when quality depends on precise, located information — internal procedures, product documentation — and you have to cite your sources rather than generate as best you can. In both cases, the compliance load goes up: this is not decided for convenience.
Which guardrails should be in place before authorizing actions (write, publish, delete)?
Apply least privilege, with read-only by default and write functions declared as such. Require human validation on destructive actions and strict stop criteria: action budget, time, cost, evidence. Make every operation idempotent so that a replay does not create a duplicate. Systematically log who asked for what, what the agent did, and on which rule it did it.
How do you log and audit an agent (prompts, decisions, actions) for debugging and compliance?
Log in a structured format: inputs, outputs, prompt version fingerprint, tool calls, errors, latency and run identifier. Also keep the decisions — stop, escalation, refused action — with the rule that fired, and the artefacts produced. The sufficiency test is simple: you have to be able to replay a complete run from the logs alone, without reopening the original systems.
How do you size the costs and limit needless iterations?
Set budgets — number of calls, context size, duration — cache the retrievals, and replace re-generation with deterministic checks. Measure the cost per run and cut the loops beyond a defined threshold, logged when it fires. The most profitable lever remains replacing a question put to the model with a check written in code: it is faster, cheaper and stable.
Which tests should be automated to make an agent reliable before production?
Automate unit tests on the tools, integration tests on the exchange contracts, and reference outputs with a controlled tolerance. Add security tests — permissions, secrets, file paths — and business scenarios that validate concrete acceptance criteria. Finally, check non-regression: same input, same decision, within an acceptable range. A correct refusal is a test result just as much as a success is.
Continue reading
- Your project holds and your tests pass: now they have to be replayed at every change, or you have to judge an agent project somebody is proposing — that is the ground of the GitHub AI agent.
- You have picked a library and have to check what its licence allows, especially if the agent goes out inside a product: licences, sovereignty and operating load are compared on the open source AI agent.
- Your tools have to reach the company’s real systems: service accounts, permissions, authorized sources and separate environments belong to AI agent integration.
- You do not want to write everything and are looking for what already exists: models, vendors and automation tools are compared on an AI agent platform.
.png)
%2520-%2520blue.jpeg)

.jpeg)
.jpeg)
.avif)