Tech for Retail 2025 Workshop: From SEO to GEO – Gaining Visibility in the Era of Generative Engines

Back to blog

AI Agents on GitHub: Judging a Repository and Framing Its Execution

GEO

Discover Incremys

The 360° Next Gen SEO Platform

Request a demo
Last updated on

26/9/2026

Chapter 01

Example H2
Example H3
Example H4
Example H5
Example H6

You are looking at a repository you did not write, and the question is not whether it is impressive: it is whether you can depend on it next month. Or else it is your own agent that has to run every morning without anyone watching it start. In both cases, what decides is not the power of the model, but what the forge lets you prove: who changed what, what was tested, what was authorized to run. A GitHub AI agent is judged on public signals, then run inside a chain where every step produces a verifiable output. For the general method — framing, level of autonomy, specification, acceptance testing — knowing how to create an AI agent gives the frame within which everything below applies.

 

What a code forge brings to an agent: versioning, testing, auditing

 

An agent is not an ordinary program. Its behaviour depends on instructions written in natural language, on data that moves, and on a model that does not return exactly the same output twice. A forge fixes none of those three things: it does something else, and that is what makes it decisive. It makes the agent’s actions versioned, tested, audited. Three benefits follow.

  • Speed: iterating quickly on prompts, templates, extractors and validators, without starting again from a copy of a file sent in a message.
  • Governance: logging “who changes what” through reviews, continuous integration checks and branch protections.
  • Reuse: pooling building blocks between teams and between countries, instead of leaving everyone to rewrite the same connector.

Those three benefits assume repository conventions, and those are set before the first automated contribution. Code assistance is positioned as an accelerator of implementation and standardization, not as an autonomous decision-maker. So impose a style and a project structure — static analysis, automatic formatting, standard folders. Add unit tests and integration tests around the agent’s “actions”, that is around everything that writes, publishes or calls a third-party system. Document the inputs, the outputs, the limits and the failure scenarios. And in review, check above all the handling of secrets, access rights and sensitive data: that is where the errors no test catches take up residence.

Then comes the main trap, and it is not technical. An agent can “look” good in a demo, then drift in production because of non-determinism, incomplete data or a badly controlled context. Your guardrail is a versioned pipeline: prompts — or instructions — in the repository, tests, evaluation sets, and usable logs. Add to that the dependencies — packages, models, third-party interfaces — which age faster than your code, and compliance, which forces you to write down in black and white what the agent does alone and where human validation stays mandatory. In practice: treat your prompts as code, and require a review before any change that affects public content.

 

Reading an agent repository: what the signals say, and what they do not

 

A repository’s front page is a shop window, and it is optimized as one. Stars give a signal of popularity, but not of production quality: they say a project was noticed, often on the day it was announced, never that it holds for three months in an execution chain. Forks indicate active reuse, issues reveal the friction zones, and pull requests show the contribution dynamic. None of these signals is read on its own: each raises a question whose answer is elsewhere in the repository.

 

Four signals, and what has to be checked behind them

 

Four entry points are enough to cover a project in ten minutes, and they are read in the order of what commits you most: release cadence, because it governs your own version bumps, actual activity, how the project handles what does not work, and how it accepts other people’s work.

Signal What it indicates What should raise a flag What you have to check
Releases Release cadence No version for months, or versions with no notes Changelog, compatibilities, migrations
Recent activity Real maintenance Frequent contributions that only touch the documentation Last update, responses to issues
Issues Perceived quality and bugs Threads open for a long time where nobody answers any more Types of bugs, resolution times, duplicates
Pull requests Community governance Contributions merged with no review and no tests Review, continuous integration tests, quality of the discussions

 

The most misleading signal is the second: sustained activity can be no more than a refresh of the shop window. Open the latest changes and look at what they touch. A living project fixes behaviour, closes tickets and publishes migration notes; a project at the end of its life updates its front page file.

 

Five families of project behind the same word

 

Before even judging a repository’s quality, you have to know what you are looking at: on a forge, “agent” can mean very different things, and half the disappointments come from there. Five families cover most of what you will meet.

  • Orchestration framework: a library with which you write your own agent. You will depend on it for a long time, so its life cycle commits you.
  • Ready-made agent: a command-line tool or a full application. Quick to try, hard to adapt.
  • Infrastructure building block: execution sandbox, memory, browser automation. These are the most sensitive on the security side, because they execute or store.
  • Multi-agent and role orchestration: dividing the work between several agents. Attractive in a demo, expensive to run.
  • Teaching resource: examples, tutorials, link collections. Very useful for understanding, never intended for production.

The most frequent confusion concerns the last family: a widely followed teaching resource is often the first result you land on, and it was never written to run in production. Classify the project before assessing it; the criteria that follow do not apply the same way to a library and to a tutorial.

 

Deciding: depend on it, fork it, or take inspiration from it

 

For enterprise use, the sorting must not stop at popularity. You are looking for a maintained, auditable, testable project with a manageable risk surface. Five criteria are enough, and the first can end the examination on its own.

  • Licence: compatible with your internal policy and your customer constraints. What that text really allows, down to the model weights and redistribution, is handled on the open source AI agent.
  • Maintenance: recent activity, releases, responses to issues.
  • Security surface: code execution, network calls, data storage, secret handling.
  • Evaluation: presence of tests, reproducible examples, datasets, metrics on quality, latency and cost.
  • Governance: contribution rules, continuous integration, code review, transparency on the roadmap.

 

What you read first, and how you know maintenance is real

 

Reading order matters, because it makes you drop early the projects that will not pass. The licence first, ten seconds, and the examination stops if it is incompatible. Then the front page file: look in it for what it does not say — the known limits, the failure cases, the permissions required. Then the tests: whether they exist, but above all what they cover. A project that tests the formatting of its outputs and not its tool calls will leave you to discover the real problems in production. Finally the examples: the one you can replay as it stands is worth more than a page of description.

Real maintenance is checked the same way: by looking at what happens when something does not work. Open three recently closed tickets and read the discussion. A maintainer who asks for a way to reproduce, fixes it, then publishes a version, is maintaining; a repository where tickets close through inactivity is not. And if you are assessing the structure and the tests because you are wondering whether you would not write the loop better yourself, then it is a Python AI agent you are about to build, and that is no longer the same decision.

 

The three possible outcomes, and what each one commits

 

An examination that does not end in a written decision will be redone in six months, by somebody else, with the same doubts. Three outcomes only, and each costs in a different place.

  • Depend on it: you freeze a version, you follow the security advisories, and you accept living with the project’s rhythm. Reserved for repositories that pass the five criteria, and governance above all.
  • Fork it: you take control of the code, and you also take on the load of maintaining it. It is a team decision, not a developer’s: it commits somebody’s time, every month, with no end date.
  • Take inspiration from it: you read, you take the ideas, you write your own. Often the right answer for a narrow building block, almost always when faced with a teaching resource.

The trap is leaving the decision by default: nobody settled it, the project simply arrived among the dependencies through one quick first integration. Write the decision and its date into the repository that uses it, with the two or three points that motivated it: when the project is abandoned, that note will tell you in a minute whether to migrate or to take over maintenance.

 

Running an agent in an integration chain

 

An integration chain automates complete sequences: build, tests, deployment, but also scheduled execution of agentic tasks — collection, enrichment, generation, checking. The critical point is not “can it be automated?” but “can it be automated while keeping control: permissions, evidence, rollback”. An integration chain becomes necessary precisely there: when what runs has to be tested, versioned and audited.

 

The six objects of an execution chain

 

Six objects are enough to describe any sequence, and they are all decided before the first run. The first three say what runs, the last three say what it produces and with which rights.

  • Triggers: launch on a contribution, on a proposed change, on a schedule, or on an external event.
  • Jobs: separate steps — extraction, generation, validation, publishing — rather than a single block impossible to resume.
  • Runners: the execution machines, hosted by the platform or by you.
  • Artefacts: storage of the outputs — reports, datasets, logs — attached to the run that produced them.
  • Secrets: keys and access tokens, to be isolated and rotated regularly.
  • Permissions: a least-privilege model, essential as soon as the agent acts and not only when it reads.

 

A run laid out, from trigger to publishing

 

The full sequence fits in five stages, and each is judged on what it leaves behind. The trigger records the event, the repository version used and the settings chosen: without that, you will never replay a run. Extraction produces a dated dataset kept as an artefact, because that is what will be compared when an output looks abnormal. Generation calls the model and writes in a constrained format, never directly into the destination system. Validation puts those outputs through the automated checks and refuses whatever does not satisfy them. Publishing only happens if the previous stage passed, and it records the state before and after.

What stops the chain is decided at the same moment, not during an incident: a step that fails stops cleanly and goes back to a human with its context — the artefact, the log, the input that caused the problem — rather than restarting blindly. Inputs and outputs are then handled as contracts: an interface or an external event can trigger the chain, call quotas are managed explicitly, observability rests on structured logs, artefacts, traces and metrics, and an alert fires on drift — cost, latency, failure rate, abnormal volume. A run that succeeds without producing evidence is not a successful run: it is a run you will know nothing about the day it fails.

 

Patterns that do not break, and what blocks a merge

 

A well-built chain looks like an assembly line: every step produces a verifiable output, and the whole can resume without everything breaking. Four patterns are enough for that, and they are set in the first version because they are almost impossible to add later.

  • Scheduling: running monthly refreshes and daily quality checks without overloading the chain or the quotas.
  • Validations: blocking publication if a test fails — missing sources, invalid structure, duplication.
  • Rollback: going back to the previous version as soon as a change degrades the result, without rebuilding by hand.
  • Idempotency: rerunning without producing duplicates or side effects.

Then comes the most concrete question: what exactly stops a bad output reaching production? Split your checks into two categories, and write that split down somewhere. Blocking are the checks whose failure signals damage: a secret detected in a change, a template that does not render, a missing source, a non-regression test that falls, a quality rule violated. Informational are those that flag a deviation with no proof of damage: test coverage falling, latency rising, a dependency warning. The temptation is to make everything blocking; that is the surest way to get a team that overrides them by reflex.

So plan for the override, instead of pretending it will not exist. One identified person may force a merge, never the author of the change, and never without a written reason. That override is logged like everything else: who, when, which check, why. Reread that list once a quarter: it will tell you either that a check is badly calibrated and should be demoted to informational, or that a team has picked up a habit that has to stop.

 

Securing execution

 

A workflow that runs an agent is an attack surface, and more so than an ordinary workflow for one precise reason: part of what it runs comes from the context it retrieved itself. Booby-trapped content, picked up by an extraction step, can influence what the next step asks the model. This risk is not handled through vigilance: it is handled through rights.

Apply least privilege to the permissions: a chain that produces a report does not need to write to the repository, and a chain that opens a proposed change does not need to merge it. Limit access to secrets by environment, with distinct values in development, in staging and in production, and rotate them regularly rather than after an incident. Prefer isolated runners as soon as the agent runs generated code: at that moment, you are no longer running your code, you are running a model’s, and the machine doing it must carry nothing else. Finally, check that secrets cannot come back out: in a log, in a kept artefact, in a copied error message.

Then come third-party actions, that is the code you did not write and which nevertheless runs with your rights. Review them the way you review a software dependency: frozen version, provenance, and regular audit. An action referenced by a moving name can change its content without anything changing on your side; that is exactly the scenario a frozen version removes. Keep the list of authorized actions, revise it on a fixed date, and treat an unknown maintainer as a reason to refuse.

 

Versioning everything that changes behaviour

 

If you use instructions or templates to have a model produce something, version them like code. This is not an analogy: a changed prompt changes the system’s behaviour exactly as a changed line of code does, except that nothing tells you and no compiler objects. The practical rule: if an object can change the output without any code moving, it belongs in the repository and it goes through review.

 

The four objects to put under version control

 

Four families of object cover the essentials, and each calls for a different check: what protects a prompt does not protect a dataset. The table below gives, for each, the recommended check and what breaks when it does not exist.

Versioned object Why Recommended check What breaks without it
Prompts and instructions Avoiding drift in tone and content Review + tests on a sample Nobody knows which wording produced which output
Sources and datasets Tracing freshness and origin URL check + date A drop in quality becomes impossible to attribute
Output templates Standardizing the structure of the deliverables Automated validation + rendering Every run returns a result in a different shape
Quality rules Automating the guardrails Blocking CI checks The guardrail goes back to being a matter of human vigilance

 

Testing a change before merging it

 

Versioning without testing merely documents the drift. So build a representative test set — simple cases, edge cases, expected failures — and measure three dimensions every time: output quality, latency and cost. The last two are the ones that get forgotten, and they are what decides what you can afford to run every day. That test set is built once and added to after every incident: a case that cost you an evening goes in the next day, otherwise it will come back.

Then add non-regression tests at every change of prompt or template. The criterion is simple: on the same input, the output has to stay compliant with the rules, even if its wording varies. Comparing two texts word for word makes no sense with a probabilistic model; checking that an expected structure is present, that a source is cited and that no rule is violated does. Finally, keep the artefacts of every run for audit and comparison: that is what will let you say, in three months, which version produced which result.

 

FAQ on AI agents on GitHub

 

Which agents are available on GitHub (open source repositories, frameworks worth knowing, examples)?

 

Five families stand out, and naming them is worth more than counting them: orchestration frameworks with which you write your own agent, ready-made agents as a command-line tool or an application, infrastructure building blocks — execution sandbox, memory, browser automation — multi-agent projects that divide the roles, and teaching resources. The first three are meant for production, the last never.

 

How do you quickly identify a reliable AI agent repository (activity, governance, security, licence)?

 

Check the licence first, then recent activity, the existence of published releases, the quality of the tickets and proposed changes, and the presence of tests. Then audit the dependencies and see whether the repository clearly documents its limits, the permissions required and the handling of secrets. A project that says nothing about its failure cases has not met any, or does not publish them: both are a problem.

 

What is GitHub Copilot?

 

It is development assistance: it proposes code and completions inside the forge’s environment and in compatible editors, from the project context. It speeds up implementation and standardization, but it does away with neither tests, nor review, nor security checks. Rolling it out across an organization — licences, rights, policies — is a decision separate from individual use.

 

How do you use code assistance with GitHub for an agent (quality, tests, review)?

 

Use it to produce skeletons — connectors, parsers, interface wrappers — then lock quality down through tests and continuous integration. Require systematic review, demand tests on the critical actions (writing, publishing, data access) and document the failure scenarios. The rule that holds: assistance speeds up implementation, it decides nothing.

 

How do you automate with GitHub (GitHub Actions, CI/CD and workflow automation)?

 

You define triggers — contribution, proposed change, schedule — jobs, runners, and strict handling of secrets and permissions. The reliable path is a sequence in steps: extraction, generation, validation, publishing, with an artefact kept at every step for audit, then an alert if the chain drifts in cost, in latency or in failure rate.

 

What is the difference between an “agent” published on GitHub and a GitHub Actions workflow enriched with AI?

 

A published agent is a project, that is code describing a logic of action: tools, memory, planning. An execution chain is an orchestration mechanism: when, how, with which permissions — and you can connect an agent to it. One provides a governed execution frame, the other is the intelligent building block. Confusing them means looking for governance in the wrong object.

 

Which good practices prevent secret leaks and uncontrolled execution in GitHub Actions?

 

  • Apply least privilege to permissions and tokens, in write mode as in read mode.
  • Segment by environment — development, staging, production — with separate secrets.
  • Switch on secret rotation, limit their validity in time, and check that they appear neither in a log nor in an artefact.
  • Isolate the runners as soon as generated code runs, and audit third-party actions at a frozen version.

 

How do you test and evaluate an agent (quality, cost, latency) before opening it to a team?

 

Build a representative test set — simple cases, edge cases, failures — and measure three dimensions: output quality, latency and cost. Add non-regression tests at every change of prompt or template, and keep the artefacts for audit and comparison. Then open it to a small scope before the whole team: that is where the uses you had not thought of show up.

 

Continue reading

 

  • The problem is upstream of the remote repository, on the workstation: running a session, setting what is forbidden on an open repository and reviewing the diff belong to the VS Code AI agent.
  • Your chain runs, but the business sequence itself still has to be drawn: steps, branches, conditions and cost per deliverable are handled on the AI workflow agent.
  • Your agent has to reach the company’s systems: connection modes, service accounts, authorized sources and logs are the subject of AI agent integration.
  • You are looking for which projects and tools exist before judging one of them: models, vendors and automation tools are compared on an AI agent platform.

Discover other items

See all

Next-Gen GEO/SEO starts here

Complete the form so we can contact you.

The new generation of SEO
is on!

Thank you for your request, we will get back to you as soon as possible.

Oops! Something went wrong while submitting the form.