How to Test an AI System Before Customers Do

  • Home
  • <
  • Blog
  • <
  • How to Test an AI System Before Customers Do
A developer reviewing code and system flow diagrams across multiple screens
AI QUALITY & EVALUATION

How to Test an AI System Before Customers Do

Neo Hives IT Solutions· 10 September 2026·13 min read

"It seemed fine when we tried it" is how almost every AI quality problem reaches a customer. The demonstration is the least reliable test in existence, because the person running it asks the questions they already know it handles, in the phrasing it likes, and stops when it works. Nobody does this dishonestly. It is simply what happens when there is no scorecard.

This is a guide to AI evaluation — how to build a test set for a system whose output is different every time, how to score it without hiring a research team, what to test for beyond correctness, and how to keep testing after launch. It is the least glamorous part of this field and the strongest predictor of whether a project survives its first year.

Why normal software testing does not transfer

Conventional tests assert equality: this input produces exactly that output, or the build fails. Language models break that contract in three ways at once. The same input can produce different wording each time. There is often no single correct answer — a summary can be right in twenty phrasings. And quality is not a boolean but a distribution: the honest statement about a classifier is "96% correct on a representative sample", not "it works".

So the unit of testing changes. Instead of assertions that pass or fail, you need a fixed set of examples with known good answers, and a scoring method — an evaluation set, run on every meaningful change, producing a number you can compare with last week's. That is the entire discipline. Everything below is detail on doing it cheaply.

Building an evaluation set in a day

The instinct is to write examples from imagination. Do not: you will write the cases you already thought about, which are the ones the system already handles. Take them from reality — past support tickets, real invoices, actual call recordings, the last two hundred enquiries. A hundred real examples beats a thousand invented ones, and a hundred is genuinely achievable in a day.

Composition matters more than size. A set made entirely of straightforward cases will report 98% and tell you nothing. Aim roughly for this shape:

CategoryShareWhat it catches
Ordinary cases~40%Baseline competence, and regressions in the common path
Edge cases~30%Unusual formats, missing fields, multiple languages, very long or very short inputs
Should refuse or abstain~10%Questions the data cannot answer, requests outside scope — the category that exposes guessing
Adversarial~10%Prompt injection, attempts to extract data, deliberately misleading phrasing
Past bugs~10%Every issue you have ever fixed, so it cannot come back quietly

That last row is the highest-value habit in the whole practice. Every time something goes wrong in production, the fix is not complete until the case is in the set. This is what stops a system oscillating — solving today's complaint by reintroducing last month's.

One discipline to keep: do not tune your prompts against the same examples you score on. If you iterate until the set passes, the set has stopped measuring anything — it has become part of the prompt. Keep a development split for iterating and a held-back split you look at less often.

How to score without a research team

Use the cheapest method that genuinely detects failure, in this order:

  • Deterministic checks first. Does the JSON parse? Are all required fields present? Is the total the sum of the lines? Is the invoice number a real invoice number in your database? Does the cited document actually exist? A surprising share of production failures are caught here, for free, with no judgement involved.
  • Exact or near-exact match where there is one right answer — extraction, classification, routing. These are the workloads with clean numbers, and they are also the ones businesses most often deploy, so this covers more ground than people expect.
  • Human rating with a binary rubric for open-ended output. Not a five-point scale, which two reviewers will never apply the same way — a short list of yes/no criteria: is every claim supported by the source? is anything important missing? is it the right length and tone? is anything unsafe? Binary criteria are boring and they agree between raters, which is the only property that matters.
  • A model as judge, once you have validated it. Useful for scale, and it must earn trust first: have it grade a set a person has already graded, and check it agrees. Judge models are known to favour longer answers, to be swayed by confident phrasing, and to be lenient towards output from a model like themselves. Validated against human labels on your own task, they are a genuine multiplier. Unvalidated, they are an expensive way to generate reassuring numbers.

What to test for beyond "is it correct"

Failure modeHow to test itAcceptable level
Invented factsQuestions whose answers are absent from the source materialNear zero; it should abstain instead
Failure to abstainDeliberately unanswerable questionsShould abstain nearly always
Prompt injectionDocuments and emails containing instructions aimed at the modelZero successful instruction-following
Data leakageThe same questions asked as users with different access rightsZero tolerance, tested every release
Format breaksLong, empty, non-English and malformed inputsMust fail cleanly, never silently
Over-refusalPerfectly legitimate requests that sound sensitiveRare — an over-cautious system quietly destroys adoption
Latency95th percentile, not the averageSet per channel; voice is far stricter than email
Cost per taskMeasured at real volume, including retriesTracked over time, with an alert on the trend
Tone and brandA sample reviewed by whoever owns the brandConsistent; this is where a rubric beats an opinion

Two of those deserve emphasis because they are so often missed. Over-refusal is a real failure: a system that declines reasonable requests trains staff to stop using it, and nobody files a ticket saying "it was too careful", so it never appears in your metrics. And prompt injection is not theoretical once your system reads content it did not author — a supplier invoice, an inbound email, a web page can all carry text designed to redirect the model. The OWASP Top 10 for LLM applications is the sensible starting checklist, and the mitigations are architectural: never let untrusted text reach a component with authority, keep tools narrowly scoped, and require approval for anything irreversible.

Deciding what "good enough" means — before you look

Set the threshold in advance. Once results are on the screen, the number will look either encouraging or disappointing, and the temptation to reason backwards from it is overwhelming. Written down beforehand, a threshold is a decision; written afterwards, it is a rationalisation.

The threshold is not one number, either. It should be per category and tied to consequence — because the cost of being wrong is what actually varies:

  • Safety and permissions: zero failures. Not a percentage.
  • Financial or legal output: high accuracy plus mandatory human approval. Do not trade one against the other.
  • Routing and classification: whatever beats the current human error rate, which is rarely as low as people assume — measure it before you set the bar.
  • Drafting for a person to edit: a lower bar is entirely rational. The question is whether editing the draft is faster than writing from scratch.

This is also where the honest comparison lives. The benchmark is not perfection, it is your current process, measured. Teams routinely reject automation at 94% accuracy while their manual process runs at 91% and has never been counted.

The model underneath you will change

Every model you build on is eventually retired, repriced or quietly updated, and this is the most underrated operational risk in the field. Without an evaluation set, a forced migration is a period of guessing — you cannot tell whether the replacement is better or worse, so you rebuild trust from zero. With one, it is an afternoon: run the suite against the candidate, compare the numbers, adjust the prompt, ship.

  • Pin the exact model version in production. Silent upgrades are how a working system changes behaviour on a Tuesday for no reason anyone can name.
  • Keep prompts in version control with the evaluation results attached to each change. A prompt is code; treat it that way.
  • Re-run the suite when anything moves — model, prompt, retrieval settings, document corpus, tool definitions. All five change behaviour, and only two are obvious.
  • Re-qualify cheaper models periodically. Prices fall and small models improve; the workload you needed a large model for last year may not need one now. The evaluation set is what lets you find out safely instead of guessing.

Testing does not stop at launch

Pre-launch evaluation tells you the system works on cases you thought of. Production tells you what people actually do, and it is always broader. What to watch, in rough order of value:

  • Error escape rate — wrong outputs that reached a customer or a ledger. Hardest to measure, most important to know. Sample deliberately; do not wait for complaints.
  • Implicit feedback. Users editing the draft heavily, re-asking the same question, escalating to a human, or abandoning halfway. These are your real quality signals, and they are free. Explicit thumbs-up buttons are used by almost nobody.
  • Abstention and escalation rates over time. A drift here usually means your documents or your inputs changed, not the model.
  • Latency and cost per task, tracked continuously with alerts on the trend rather than reviewed quarterly.
  • A sampled human review, weighted towards low-confidence cases. Even a few examples a week catches classes of problem that no automated check anticipated.

The channel changes the emphasis. For a voice agent, latency and interruption handling matter as much as correctness. For a knowledge assistant, retrieval hit rate and citation accuracy come first — and when an answer is wrong there, it is usually retrieval at fault rather than the model, which is only diagnosable if you log what was retrieved.

Where evaluation goes wrong

  • The evaluation set has twelve examples, so any result is noise.
  • Prompts were tuned against the same examples used to score, so the number measures memorisation.
  • Only the average is reported, hiding a category that fails completely.
  • The set was changed after a bad result rather than the system.
  • A model judge was trusted without ever checking it against human ratings.
  • No baseline was taken from the manual process, so "better" is unprovable.
  • Safety and permission tests were run once, before launch, and never again.
  • Adversarial inputs were never tried, because nobody on the team thinks like an attacker.
  • Evaluation was owned by the vendor, so the client has no independent view of quality.

How we approach this at Neo Hives IT Solutions

We build the evaluation set before the system, from your real cases, and we hand it to you — it is your asset, not ours, and it outlives whoever built the first version. Thresholds get written down before results are seen, deterministic checks come before human judgement, and anything touching permissions or money is tested with zero tolerance rather than a percentage.

After launch the same suite runs on every change, with sampled review and a small set of monitored numbers rather than a dashboard nobody opens. This is the work our testing and quality checks for AI service exists to do, and it sits naturally alongside conventional software testing — the same instincts, different failure modes. It is also the honest answer to "how do we know it works", which is the question every serious buyer asks and surprisingly few vendors can answer with a number. You can see how we report outcomes in our case studies.

Common questions

How many examples do we need? Fifty is enough to be useful, a hundred to be reasonably confident, a few hundred if a percentage point matters commercially. What matters more than size is that they are real and that the awkward categories are represented.

Who decides the correct answers? The person who does the work today — and it is worth noticing how often two experienced people disagree. When they do, you have found an ambiguity in your own process, not a problem with the test. Resolve it in the specification before blaming the model.

Can we use a model to grade instead of people? Yes, after validating it against human grades on your own task, and best restricted to the criteria it agrees with humans on. Use it for volume; keep a human sample as the anchor.

How often should the suite run? On every change to prompt, model, retrieval or tools, and on a schedule regardless — weekly is plenty for most systems. The scheduled run catches drift in your inputs, which is the failure nobody triggers deliberately.

Can we just A/B test in production instead? That measures whether users prefer something, which is worth knowing but is not the same as whether it is correct. A confidently wrong answer often scores well with users. Run both; do not substitute one for the other.

Is this not expensive? It is a day or two to build and minutes to run. The comparison is not against zero cost, it is against debugging a quality complaint from a customer with no way to reproduce it — which is the most expensive way to discover any of this.

Where to start

Take twenty real cases from last month — tickets, invoices, enquiries, whatever your system will handle — and write down the correct outcome for each. That is a working evaluation set, and it will already tell you something uncomfortable about whatever you have running today. Add the awkward categories next week.

If you have an AI feature in production and cannot say what its accuracy is, that is the gap worth closing before adding the next feature. Our free AI readiness audit includes a look at how whatever you already have is being measured, and getting in touch costs nothing if you just want a second opinion on a number.