Why One Good AI Answer Isn't Proof of Anything
A single successful output from an AI system reveals almost nothing about how it will perform in the real world. Rigorous evaluation requires multiple test inputs, repeated runs, and statistical analysis.

Picture asking an AI customer-service chatbot the same question three times: "Can I return an opened product after 30 days?" The first response accurately describes the return policy. The second omits a critical exception. The third confidently guarantees a refund the customer isn't actually eligible for. Which answer represents what the system truly does? The honest answer is all three—and none of them in isolation.
Yet many teams assess AI systems by running them once, examining the output, and making judgments about capabilities and deployment. This happens not because teams are negligent, but because decades of traditional software have conditioned us to expect consistency: if a feature works once, it works the same way every time. AI systems operate under fundamentally different rules. One successful output proves the system can do something. It says nothing about whether it will do it reliably.
AI Outputs Are Nondeterministic
A nondeterministic system can produce different outputs when given the same input.
Language models work by assigning probabilities to potential next tokens and selecting from them as they build responses. Feeding the same question into the system multiple times can yield answers that vary in both wording and quality. This variation can be substantial. Studies comparing repeated outputs from identical language models have documented meaningful performance differences across runs and shown that evaluation findings can hinge on how outputs are generated.
A single output is an example, not an evaluation.
AI Evaluation Should Resemble a Quantitative UX Study
Consider how researchers would approach a quantitative usability study of an e-commerce checkout process. They would not have one person complete one checkout and declare the site perfectly usable because that person succeeded. Instead, they would design a collection of representative checkout tasks involving different product types, observe many participants attempting them, and report metrics like task success and time spent, along with confidence intervals showing how much those estimates might vary across the broader user population.
AI evaluation follows the same logic. Rather than asking people to interact with an interface, the evaluator asks an AI system to generate outputs for a set of test inputs.
The parallel is not perfect—an AI run is not equivalent to a human participant. Yet both approaches depend on collecting multiple observations to estimate system performance. Quantitative UX researchers synthesize results across tasks and participants instead of relying on one successful attempt. AI evaluations should do the same: use representative inputs, gather sufficient observations, and report average performance alongside variability and uncertainty.
A single output demonstrates only that the system succeeded once. A proper evaluation submits many representative inputs, runs each repeatedly, and summarizes all outputs as an average with a confidence interval. The underlying principles come from experimental design and statistics—fields with decades of established practice. AI measurement does not require inventing new science.
What Would an AI Evaluation Look Like?
Imagine evaluating an AI system designed to answer customer-service questions based on company policies. The evaluator might select 10 representative questions spanning different topics, complexity levels, and customer scenarios—perhaps covering returns, cancellations, damaged orders, subscriptions, and warranties. Each question gets submitted to the system 5 times, generating 50 total answers. (These numbers are illustrative, not universal prescriptions. A real evaluation might require more questions or runs depending on system variability and decision importance.)
Before running the evaluation, define what makes an answer acceptable. For instance, an answer might qualify as acceptable only if it:
- Correctly answers the customer's question
- Includes all relevant conditions and exceptions
- Aligns with company policies
- Tells the customer what to do next when needed
Any answer failing these criteria counts as unacceptable. (This example uses a binary quality metric, though more nuanced metrics are possible.)
The evaluator then calculates the percentage of acceptable answers for each question across its repeated runs and summarizes performance across all 10 questions. The results should include:
- The percentage of outputs deemed acceptable
- A confidence interval around that percentage
- How performance varies across different questions
- How consistently the system answers the same question
A confidence interval represents a range of plausible values for average system performance. It serves as a reminder that any single evaluation score is an estimate, not a fixed and permanent characteristic of the system.
The purpose of this example is not to mandate 10 questions and 5 runs universally, but to illustrate the core structure: test multiple representative inputs, run each input repeatedly, and summarize the resulting distribution of scores.
Any evaluation represents a snapshot of a specific system at a specific moment. Like UX-benchmarking studies, all details must be carefully recorded: the model and version, prompts or instructions, settings, tools, context, and evaluation date. If any component changes, the evaluation may need repeating.
Two Questions an Evaluation Must Answer
The example above helps answer two distinct questions:
- How well does the system perform across the range of questions users may ask?
- How consistently does it answer the same question?
These questions capture two forms of data variability: test-input variability and run-to-run variability. In the customer-service scenario, differences in answer quality across the 10 questions represent test-input variability. Differences among the 5 answers generated for the same question represent run-to-run variability.
Test-Input Variability
A customer-service system might excel at simple, direct questions like "Where is my order?" but struggle with questions involving ambiguous policies or multiple conditions. It may correctly explain the standard return policy but fail when a customer's situation qualifies for an exception.
An evaluation dominated by easy questions will therefore overstate real-world performance.
Test-input variability represents variation in the outputs that results from the particular inputs included in the evaluation.
Adding more questions reduces reliance on specific examples chosen. Yet quantity alone is insufficient. A large test set containing only straightforward questions yields a precise answer to the wrong question. Test inputs must be representative of what actual users will submit in practice.
Run-to-Run Variability
Because the system is nondeterministic, it may deliver different-quality answers when the identical question is submitted repeatedly. This is run-to-run variability.
Run-to-run variability occurs when the system produces different-quality answers for the same test input.
Repeating an input reveals whether the system handles it consistently. Without repeated runs, distinguishing between a task the system performs reliably and one it completes successfully only occasionally becomes impossible.
Why You Need Both More Inputs and More Runs
More test inputs and more runs address different questions:
- More test inputs refine estimates of performance across the intended input range. In the customer-service example, they show how well the system likely handles the variety of questions customers will ask.
- More runs refine estimates of consistency for the same input. In the example, they show how consistently the system will answer the same question when different customers ask it.
One cannot substitute for the other.
The Same Overall Score Can Conceal Different Problems
Suppose two customer-service systems are each tested on 10 questions with 5 runs per question. Both produce acceptable answers on 80% of the 50 runs.
The first system answers 8 of 10 questions correctly every single run and consistently fails on the remaining 2. The second answers each question correctly in 4 out of 5 runs.
Their success rates are identical, but the systems have fundamentally different problems. The first is predictable: it consistently handles certain question types and consistently fails on others. An organization might identify those unsupported questions and route them to a human agent.
The second is unpredictable: any customer question might receive an incorrect answer. It could still be useful if outputs are reviewed before reaching customers—for instance, when the system drafts a response that a human agent checks and corrects.
The overall score alone does not reveal this distinction. Understanding an AI system requires knowing both how well it performs and how its failures distribute.
Rigorous AI Evaluation Is Not New
Some established AI evaluations already incorporate repeated attempts and account for variability and uncertainty. Coding benchmarks employ pass@k—a metric estimating the probability that at least one of k generated attempts succeeds, requiring repeated runs. Chatbot Arena publishes confidence intervals around model ratings. Anthropic researcher Evan Miller has also advocated for analyzing AI evaluations like experiments, using standard errors, paired comparisons, and methods accounting for related test questions.
Rigorous methods exist. In practice, they concentrate in academic research papers. Product teams do not necessarily need to adopt the exact metrics from academic benchmarks. However, they should apply the same underlying principles: collect multiple observations, account for important sources of variation, and communicate uncertainty around results. In other words, do not treat one output as sufficient evidence.
The Purpose of the Evaluation Does Not Change the Method
Teams evaluate AI systems for various reasons: deciding whether to adopt a new system, assessing a new AI feature's performance, or conducting quality assurance to verify a feature is ready for release. These purposes differ in what gets evaluated and why. The method, however, remains constant: representative inputs, repeated runs, and averages with confidence intervals.
Quality assurance warrants special attention. Traditional software testing assumes determinism: a passing test once is expected to pass always. For nondeterministic systems, a single passing test is a sample, not proof. A feature succeeding in a demo may still fail for 1 in 5 customers. AI quality assurance should track pass rates across repeated runs, not the outcome of one test.
Are Multiple Runs Worth the Cost?
Repeated runs demand time and money. Not every interaction with an AI system requires formal evaluation. A single run suffices for early exploration or a basic check that the system functions, but it should not serve as evidence for consequential decisions.
When an evaluation will drive a decision—launching a product feature, selecting a vendor, recommending a process, or claiming improvement—it should rest on repeated runs with multiple representative inputs. It should report average performance, a confidence interval, and information about result consistency. If repeating runs is too expensive, the honest conclusion is that evidence is insufficient.
Conclusion
When evaluating an AI system, do not ask only whether it produced a good output. Ask how often it produces good outputs across the range of inputs users will submit and across repeated runs of the same input.
Sound AI evaluation should include: (1) multiple, representative inputs, because performance on easy inputs reveals little about performance on difficult ones; (2) repeated runs on each input, because only repetition distinguishes between a system that performs reliably and one that succeeds occasionally; and (3) averages and confidence intervals, to communicate performance and uncertainty.
New statistical methods for AI evaluations are unnecessary; the same methods experimental scientists and quantitative UX researchers have employed for decades work perfectly. What is needed is stopping the practice of treating one impressive output as evidence. A single good output resembles a participant completing a task: encouraging, but not an evaluation.


