26/9/2026
Auditing an LLM: two questions that often get confused
A team puts a model on the table and asks whether it can be plugged into production. The question sounds simple; it is not. The problem often comes from an ambiguity: are we talking about auditing brand visibility in generative engines (GEO), or auditing the AI model itself (quality, security, compliance)? Both exercises carry the same name and do not answer the same person. This page deals with the second, in a deliberately restricted scope: the business assessment of a language model before adopting it — accuracy, faithfulness to sources, coverage of the need, stability of the answers — with which test set and which acceptance criteria. Security, confidentiality and bias belong to a separate assessment, which calls for other skills and other protocols: it is not opened here. The terms that recur here — hallucination, context window, benchmark, inference — are defined in the generative AI glossary.
Assessing a model, or measuring what it says about you
In practice, the phrase “auditing an LLM” covers two distinct exercises:
- Visibility audit in AI answers (GEO): measuring whether your brand, your offers and your content are mentioned, recommended and cited with sources in generated answers.
- Language model audit: assessing the AI itself (factual precision, robustness, bias, confidentiality, security, traceability). This exercise belongs more to governance and risk management, but it has an indirect marketing impact: an unreliable AI can distort your positioning and amplify errors.
The first belongs to the AI GEO audit, which measures the brand's presence in generative answers: share of voice, sources cited, accuracy of the mentions. The second does not look at what an answer says about you. It looks at what the model can do on a task whose expected result you already know. Confusing the two leads you to measure a brand when you meant to judge a supplier.
What can be checked on a model, and what depends on its version
A model is examined on observable outputs, reproducible under equal conditions:
- Accuracy on a question whose right answer is known in advance and documented on your side.
- Gap between two runs of the same question, with identical parameters.
- Compliance with the requested format: length, structure, expected fields, style instructions.
- Behaviour outside the scope: explicit refusal, invention, or a partial answer flagged as such.
- Observable constraints: language, volume of text accepted in one go, response time.
Conversely, some factors vary with the context: model, version, localization, “persona”, conversation history, and whether web search is switched on. The object being assessed still has to be named precisely: a model carries a version identifier, a consumer application wraps it in a product, and what gets plugged into production is most often a complete system — a model, document retrieval and instructions. Changing one of the three changes the outputs without the model having moved. That calls for a sampling protocol, re-tests, and tracking over time rather than a single photograph. It is this instability that governs everything else: a test set replayed identically is worth more than a conviction formed in one demonstration.
The criteria a model is judged on
If your objective is to assess a model and not only your visibility, the criteria frequently cited fall into two families that are better kept apart. The business assessment, the only one dealt with here, covers factual accuracy, faithfulness to the documents supplied, coverage of the need and stability of the answers: it decides the confidence placed in an output before review, and therefore what can be published without being reworked. Risk assessment — bias and fairness, data confidentiality, security — answers other questions and is conducted separately, with other skills; this page does not open it. That leaves traceability and usage rules, dealt with further down because they govern the reproducibility of the assessment itself. Three criteria are settled before any decision, because they rule a model out without discussion: accuracy, repeatability and usage constraints.
Factual accuracy and hallucination
A hallucination is not a clumsy turn of phrase: it is a false statement delivered with the same assurance as a true one, and therefore invisible to anyone who does not already know the answer. Two measures hide behind it, which are worth scoring separately: factual accuracy — the answer is true — and faithfulness to the documents supplied — the answer reflects the source it was given. An answer perfectly faithful to a false or outdated source is still false; a true answer can depart from the document it was meant to render. The share of users who have encountered at least one AI hallucination is 17% (Exploding Topics, 2026). The language model statistics give this benchmark with its source and its year. What gets measured is therefore not the absence of error, which no provider guarantees, but the error rate on the information that commits the company: the scope of an offer, conditions, regulatory figures, product versions. The threshold beyond which a model is ruled out is set before the run, on a named scope, and it is not the same for an internal note as for a published page.
Robustness and repeatability: the same question, twice
A model can be accurate once and wrong the next time on the same question. Repeatability is therefore measured for itself: each question is replayed several times, with identical parameters, and the gap between the outputs is noted — factual contradiction, an answer that is sometimes complete and sometimes truncated, a format that varies. Robustness adds resistance to wording: the same request put in two different turns of phrase, or in a different language, must produce the same substance. A model that performs very well on average but is unstable from one run to the next costs more than a slightly less precise but predictable model, because instability is paid for in systematic review.
Usage constraints: context window, language, cost per call
These are the criteria that rule a model out in practice, often before the question of quality. The context window sets the volume of text processed in one go: it decides whether a product repository, a corpus of pages or a whole specification fits into a single call, or whether it has to be split — and splitting means adding machinery to maintain. The working language, the expected register and the ability to hold in-house terminology are checked on your own texts, not on a demonstration in English. Cost per call and response time, finally, are related to the volume actually produced each month. The table below connects each criterion to what is observed and to what you decide from it.
Reading a benchmark without drawing the wrong conclusion
It is generally the first thing people look at, and the most badly read. A public ranking answers a precise question — which model performs best on a given standardized test — and that question is almost never yours. It serves to rule out candidates that are obviously out of the running and to spot the families of models that are improving; it does not serve to decide between two finalists for a given use.
What a score measures, and on which set
A benchmark is a closed set of tests: a set of questions, a scoring method, a way of counting the right answers. The number of main benchmarks for LLM evaluation is 6 to 15 depending on the platform (Palmer Consulting, llm-stats.com, 2026), which is enough to explain why two public rankings do not give the same order: they do not compose the same test and do not weight it the same way. The leading scores on the GPQA, MMLU, MMMU and AIME 2025 benchmarks are above 0.85 (llm-stats.com, 2026): at the top of the ranking, the gaps narrow, and a table discriminates less and less as it approaches its ceiling.
Why a public ranking does not replace your test set
A public benchmark measures a general competence on public content. Your use covers a narrow domain, an in-house vocabulary, documents the model has never seen, and an output format of your own. A model that excels at mathematical reasoning can get the scope of an offer wrong; a model that is average in the ranking can hold a structured writing instruction perfectly. A mechanical bias is added on top: public tests circulate, and nothing guarantees that a widely distributed set of questions stays unseen for a model trained after its publication. The ranking therefore serves to draw up a shortlist. The decision is taken on a test set that you own and that nobody else has run.
Building your own test set
The methods generally rely on scenario testing (sets of questions), red teaming approaches, consistency measurements (repeatability), and source checking when the model provides sources. These approaches are described here so that you know what your teams or your suppliers have to put in place: the first is within reach of any editorial team, the others belong to internal governance and security. It is the first that produces the adoption decision, and it is built on material you already have: your documents.
The questions whose answer you already know
The principle comes down to one sentence: a model is assessed only on questions whose right answer is established elsewhere and verifiable without discussion. They are drawn from your own documents — product pages, contractual conditions, technical documentation, reference pages, internal procedures — and for each one you write the expected answer before launching anything. The set covers four families: the simple facts the model has to restate without error; the edge cases where the right answer is “it depends”, with the condition; the traps, that is, questions whose answer exists nowhere and where the model has to refuse rather than invent; and format tasks, where the structure is judged as much as the substance. This set of questions is not the one used to record what an answer says about your brand: two objects, two grids.
The protocol: logged conditions, repetitions, acceptance criteria
A test set without a protocol proves nothing. Systematically document the collection conditions: date, language, persona, web search mode (on or off), and surface used. This traceability avoids confusing model noise with real progress (or real drift). Each question is replayed several times, and the scoring stays stable from one campaign to the next. Three lines are enough for the acceptance grid: Accuracy — truth of the fact and faithfulness to the document supplied, scored separately — on sensitive information, Proof (data, studies, verifiable elements) and Freshness (currency, updating). The decisive point is the order of operations: the pass threshold for each line is written before the run. Set afterwards, it always adjusts to the result obtained, and the evaluation then only serves to justify a choice already made.
Traceability, data and usage rules
An audit trail corresponds to the traceability of prompts, answers, model versions, sources (when available) and feedback. It serves to understand “who asked what, when, and on what basis”. Without it, an evaluation is not replayable and an incident cannot be explained.
Audit logs allow four things: reproducing an incident, documenting a drift (e.g. a risky answer), proving test conditions, and feeding continuous improvement — fixing a prompt, a source, a reference page, or an internal rule. These four uses are enough to decide what has to be kept: the request, the raw output, the model version identifier, the timestamp, the call parameters, and the human verdict where there was one. Anything that serves none of these uses does not have to be kept.
Four rules frame the set-up, and they fall to the company using the model: define roles and access rights, set a retention policy, avoid logging unnecessary personal data, and apply GDPR requirements to your context (without confusing a marketing article with legal advice). The objective is risk reduction, not over-collection. A policy that keeps everything indefinitely creates the exposure it claims to control.
Deciding: adopt, frame, rule out
A test set is only worth something if it leads to a written decision, and there are only three possible ones. The proportion of users satisfied with the relevance of LLM answers is 83% (Exploding Topics, 2026): perceived reliability is high, which makes the residual gap more costly, not less — a set-up that gives satisfaction most of the time disarms vigilance exactly where it would be useful.
The evaluation effort itself is sized before it is priced. There is no reliable “standard” rate without inventing figures. The variation factors, on the other hand, are stable: the number of models compared, the volume of the test set and its number of repetitions, the depth of the factual check — reviewing an answer and tracing it back to the source document do not take the same time — and the level of traceability required. A one-off evaluation before signing and a set-up replayed at every version are not organized in the same way.
What tips the decision
The three outcomes are conditioned in advance, on the grid written before the run. Adopt: the accuracy threshold is met on committing information, the outputs are stable from one run to the next, and the usage constraints — window, language, cost — cover the planned volume. Adopt with a framework: the model fails on an identified family of questions, or varies too much to be published without review. The framework is then described precisely — the scope of permitted subjects, mandatory review before publication, a ban on regulated content, data prohibited as input, logging of every call — and is re-examined on a set date. Rule out: an error on committing information, a refusal to state the model's usage conditions, or constraints incompatible with the real volume. A model that varies heavily, cites little, or distorts facts increases your reputational risk: that is a ground for rejection, not a point to watch.
What stays human once the model is in place
Collection, running the test set and comparing the outputs are automated without difficulty. The verdict on an ambiguous answer is not automated: deciding that a wording is false rather than imprecise assumes you know the file. Three control points therefore stay with the producer. On the input side, someone validates the sources supplied to the model and prohibits the data that has no business being there. On the output side, every high-stakes piece of content goes through a factual review targeted on the points the test set identified as fragile — this is where the load concentrates, and it is planned as a production step, not as an extra. At regular intervals, a sample of already published content is checked again, because a drift sets in without a signal. The residual load falls when the test set is precise; it never drops to zero.
Re-testing when the model changes
An adoption decision has an expiry date that it does not display. AI answers vary more with the context and the versions. Hence the importance of: (1) a test protocol, (2) a sufficient sample, (3) monitoring, and (4) a reading oriented towards trends rather than absolute truth.
Four occasions call for replaying the test set without waiting for the scheduled date: a version upgrade announced by the publisher, a change of supplier or of access mode, an incident observed in production, and a change to your own reference documents, which changes the expected right answers. Outside these triggers, a quarterly rhythm suits a stable use, and a closer one for sensitive scopes.
What is compared from one run to the next is narrow, and that is what makes the comparison valid: the accuracy rate on the same set of questions, the gap between repetitions, the number of correct refusals on the trap questions, and the format failures. The test set stays frozen so that two campaigns are comparable; it is extended when the scope of use widens, and each addition is dated, without touching the core that serves as the point of comparison. Producing with a model whose acceptance criteria were set in advance, with the human control point that goes with it, is the task covered by content generation framed by AI.
FAQ: assessing a language model, its limits and its governance
What does an audit look like on the GEO side and on the governance side?
On the GEO side, the exercise measures your brand's presence in generated answers, the sources that feed them and the accuracy of what is said about you. On the governance side, it assesses the model itself: accuracy on questions whose answer you know, stability between two runs, usage constraints, security and traceability. Two objects, two grids, two different decisions.
Should each model be audited separately (ChatGPT, Gemini, Perplexity, Claude)?
Yes, and each version separately, provided you name what you are testing: ChatGPT and Perplexity are applications, not model identifiers. Each wraps one or more models, document retrieval and instructions of its own, so the object assessed is almost always a complete system — that is what has to be described in the collection conditions, version identifier included. The test set is shared — that is its whole point — but it is run system by system, under the same logged conditions, then compared line by line. A result obtained on one version does not transfer to the next: a version upgrade is a ground for re-testing, not an improvement acquired.
Which sources do models use to generate their answers?
Depending on the access mode, a model answers from what it has learned, from documents you supply, or from pages consulted at the time of the request. These three regimes do not give the same reliability, and the distinction is part of the test: it is a collection condition to log. When the model displays sources, the check covers them as much as the answer.
Can the audit and the monitoring be automated over time?
Partly. Running the test set, collecting the outputs and comparing two campaigns automate well. Interpreting an ambiguous answer — false or merely imprecise, acceptable or not in context — calls for a human review. The robust approach is hybrid: automate what can be counted, reserve judgement for the cases the counting flags.
How often should an audit be carried out?
An initial evaluation before adoption, then a rhythm suited to the use: quarterly on a stable scope, closer together on sensitive subjects. To that are added the one-off triggers — a version upgrade, a change of supplier, an incident observed in production, or a change to your reference documents, which changes the expected answers.
How much does an audit service cost, depending on scope and models covered?
There is no reliable “standard” rate without inventing figures. The variation factors, on the other hand, are stable: the number of models and versions compared, the volume of the test set and the number of repetitions, the depth of the factual check, the level of traceability required. A one-off evaluation and a set-up replayed at every version are not sized in the same way.
What are logs and traceability for in a compliance exercise?
For reproducing an incident, documenting a drift, proving the conditions under which a test was run, and feeding the corrections. In practical terms, they keep the request, the raw output, the model version, the timestamp and the call parameters. The counterpart is an explicit retention policy and the exclusion of unnecessary personal data: the objective is risk reduction, not over-collection.
How often should the analysis be rerun and the prompts updated?
Set the rhythm by the stability of your market: quarterly on stable scopes, monthly on fast-moving markets. Also enrich the prompt library as soon as a new offer, a change of positioning or a reputational incident occurs, so that the protocol stays representative of your commercial reality.
Continue reading
- The model is not the subject and the question is about the site itself: the SEO audit frames the scope of a diagnosis, the moment it is triggered and the way to read its deliverable.
- The model is adopted and you now want to control what agents read of your own site: the llms.txt file has its role, its limits and its level of adoption.
.png)
.jpeg)

.jpeg)
%2520-%2520blue.jpeg)
.avif)