Getting an answer from an AI model does not yet mean the system is working correctly. In classic software, we can often determine unambiguously whether a function returned the expected result. In AI systems, an answer may be fluent, logical, and convincing, and still contain errors.
That is why, along with the development of AI, a new engineering problem appears: how do we systematically measure the quality of a system whose answers are not always identical?
This is precisely the area of AI evaluation, or the evaluation of artificial intelligence systems.
And it is much broader than checking whether a chatbot “answers well.”
Testing classic software and testing AI are not the same
Let’s imagine a simple function in an application.
The user enters: 2 + 2
The system should return: 4
If it returns 5, we have a clear error.
We can prepare a test: expect(calculate("2 + 2")).toBe(4)
and each time we will get a clear result: the test passes or fails.
In AI systems, the situation looks different.
A user may ask: “Write a short reply for a customer asking about the order fulfillment date.”
The system may generate several different answers. All of them may be linguistically correct. All of them may sound professional. But one may include an incorrect date, another may be too long, a third may omit important information, and a fourth may be perfect.
So it is not enough to check whether the answer was technically generated.
You need to check whether it meets specific quality criteria.
First problem: a good answer is not always true
This is one of the most characteristic features of generative AI.
A model can generate an answer that sounds very convincing, but is not supported by the source data.
In the case of a system using RAG, the problem is even more interesting. The system may receive a question, search several documentation fragments, and then generate an answer.
Then at least three things need to be checked:
- Were the right information found? Did the retrieval mechanism fetch fragments actually related to the question?
- Does the answer use the information found? Did the model not add something that was not present in the sources?
- Does the answer actually answer the question? You can have a correctly working retrieval, but a poor final answer.
That is why RAG evaluation separates, among other things, aspects such as the relevance of the retrieved context, retrieval completeness, answer correctness, and consistency with the sources.
This is an important shift in how we think about testing.
We no longer test only: question → answer
but the entire chain: question → retrieval → context → model → answer
You can have a good model and a bad AI system
This is another thing that is easy to forget.
A company may choose a very good language model and still create a poor AI product.
Why? Because the quality of the final system depends not only on the model.
Also important are:
- data quality,
- the way the context is prepared,
- the prompt,
- the way information is retrieved,
- model parameters,
- tools made available to the AI,
- application logic,
- memory,
- the way errors are handled,
- security safeguards,
- the way answers are evaluated.
This means that the question: “Which model is the best?”
is often less useful than: “Which model works best in our specific use case?”
A model that is excellent at generating marketing content does not have to be the best solution for document classification, data analysis, or business process support.
That is why model comparison should be done on real tasks that the system is supposed to perform.
First you need to create your own test set
It is not possible to sensibly evaluate an AI system if we do not know what we expect from it.
That is why one of the most important elements of evaluation is preparing a test dataset, meaning a set of real or representative cases.
For example, a company is building AI for customer support.
Instead of manually checking one response after every prompt change, you can prepare several hundred cases:
- simple questions,
- ambiguous questions,
- questions containing incorrect assumptions,
- questions requiring document retrieval,
- questions about exceptions,
- questions about complaints,
- questions requiring refusal,
- questions containing data that the AI should not disclose.
Each system change can then be run against the same set.
And this is exactly where AI starts to resemble classic software.
We no longer test a single response. We test the behavior of the system across the entire set of cases.
Prompts can be tested too
A prompt is often treated as text that someone wrote once and left in production.
In reality, it can be a part of application logic.
Changing one sentence can cause:
- an improvement in responses in one scenario,
- worse responses in another,
- a greater tendency to refuse,
- more hallucinations,
- longer responses,
- higher cost,
- greater token usage.
That is why a prompt should be treated similarly to code.
If we change the prompt, it is worth knowing:
- what improved?
- what got worse?
- did a regression appear?
That is precisely why automated evaluation is becoming increasingly important, rather than manually judging a few sample answers.
AI can pass the test and still be a bad product
Let’s assume we prepared 100 test cases.
The system answered correctly in 95 of them. The result looks great. But what if the five incorrect answers concerned critical situations?
If a chatbot answers questions about opening hours, five errors may be a problem.
If AI helps an employee analyze financial, medical, or legal documents, the significance of those errors may be completely different.
That is why the average alone is not enough.
We also need case weighting.
We may decide that:
- a regular question has a weight of 1,
- an important error has a weight of 5,
- a security flaw has a weight of 10,
- the disclosure of confidential information has a weight of 100.
Then the system does not get "95 percent". We get a much more useful picture of risk.
Not everything can be measured with a single number
This is one of the most important problems in AI evaluation.
We can have several metrics:
- Accuracy - is the answer correct?
- Relevance - does it answer the question?
- Faithfulness / groundedness - is it based on the provided sources?
- Context precision - are the retrieved fragments relevant?
- Context recall - did the system find the needed information?
- Safety - does it avoid undesirable actions?
- Latency - how long does the user wait?
- Cost - how much does it cost to perform the task?
So RAG can have very good answer quality, but at the same time retrieve a huge amount of context and generate an unacceptable cost.
Another system may be very cheap and fast, but make too many mistakes.
So there is no single universal number that determines whether AI is "good". Quality must be defined in the context of a specific use case.
And what about evaluating AI by other AI?
Here another interesting mechanism appears.
One way to automate evaluation is to use a model as a judge, that is, LLM-as-a-judge.
For example:
Model A generates an answer.
Model B receives the question, the answer, and specific criteria.
Then it evaluates:
- correctness,
- adherence to instructions,
- completeness,
- style,
- safety.
Modern evaluation tools also allow combining such judging with classic text comparisons, custom scripts, or graders based on specific rules.
This greatly increases the scale of testing. But it does not mean that humans are no longer needed. The evaluating model can also make mistakes. That is why in systems of greater business importance it is worth combining automated evaluations with periodic expert review.
The biggest problem: regression
Let us imagine a system that works very well. The team changes the model to a newer one. The new version is faster and cheaper. So it seems that everything is going in the right direction.
After deployment, however, it turns out that:
- the answers are less precise,
- the model refuses to answer more often,
- it uses the documentation worse,
- it interprets instructions differently,
- in some scenarios it starts providing incorrect information.
This is exactly AI regression.
In classic software we have known regression for years. In AI we also need to detect it, but the problem is more difficult because the system's behavior can change without a classic "bug". That is why every major change should be compared with the previous version.
Model.
Prompt.
Embedding.
Retriever.
Documentation.
Agent logic.
Parameters.
Each of these elements can affect the result.
An agent is not enough to evaluate just by its final answer
It gets even harder in the case of AI agents.
A classic chatbot can do one task: question → answer.
An agent can work very differently: goal → plan → tool → result → next step → decision → action → answer.
If the agent did not achieve the goal, we do not only want to know that it lost.
We want to know: where did it make a mistake?
- Did it misunderstand the task?
- Did it choose the wrong tool?
- Did it pass the wrong parameter?
- Did it retrieve the wrong data?
- Did it make a bad decision after receiving the result?
- Did it take too many steps?
- Did it stop too early?
In research on agent evaluation, more and more attention is being paid not only to the final result, but also to the execution path, tool use, planning, memory, reliability, and safety.
This means that the future of AI testing will be largely connected with analyzing traces, that is, the full execution flow of the system.
AI needs something similar to CI/CD
If AI is part of a product, it cannot be tested only before the first deployment.
The system will change.
The model will change.
The prompt will change.
The knowledge base will change.
The search method will change.
The configuration will change.
That is why evaluation should become part of the development process.
The scheme may look as follows: change → tests → evaluation → comparison with the previous version → deployment decision
If the new version improves quality in one area but exceeds the established error threshold in another, deployment may be stopped.
This is a very similar philosophy to classic CI/CD, but the criteria are different.
In the case of AI applications, we can check at the same time the answer quality, correctness, safety, cost, and latency. Research solutions are already emerging that combine evaluation with observability and quality gates in the process of deploying LLM/RAG systems.
So can AI be tested the same way as classic software?
Yes, but only partially.
The classic approach is still needed.
We test:
- API,
- integrations,
- permissions,
- data validation,
- errors,
- timeouts,
- security,
- performance,
- application logic.
But we cannot stop there.
There is a second layer: evaluation of AI behavior.
- Is the answer correct?
- Is it consistent with the sources?
- Does the model follow the instructions?
- Does the system behave correctly in unexpected situations?
- Does the agent choose the right tools?
- Has the new version not worsened quality?
- Does the operating cost remain acceptable?
- Does the user actually receive value?
This is no longer a classic unit test.
The best AI test is not always a lab test
There is one more very important element.
The system may perform great on a prepared test set, and yet still have problems in the real world. That is why it is worth observing real interactions as well. Not so that every user becomes a tester. The point is to continuously improve the system based on real cases:
- where users correct the AI,
- where they ask to repeat the answer,
- where they interrupt the conversation,
- where they escalate the issue to a human,
- where the agent does not achieve the goal,
- where unusual questions appear.
In this way, a continuous evaluation loop is created: user → AI action → result → analysis → new test case → next version of the system
This is a completely different development model than the one-time “we implemented AI and it works.”
AI should not be judged by the question “does it work?”
That is too little.
Better questions are:
- How often does it work correctly?
- In what situations does it make mistakes?
- How serious are those errors?
- Is the new version better than the previous one?
- Is the system sufficiently safe?
- Are the answers based on the right data?
- How much does it cost to achieve a specific result?
Only this kind of question set makes it possible to talk about a mature AI system.
The most important change in thinking
For years in software development, a simple rule applied: code should work.
In AI systems, this has to be expanded: the system should work well, predictably, and measurably.
That is a huge difference. Because AI is not a function that always returns the same result. It is a probabilistic system whose behavior depends on the model, data, context, instructions, and the entire architecture around it.
That is why a professional AI deployment does not end the moment the model starts answering.
That is when the real question begins: how do we know we can trust it?
And that is exactly the question a well-designed evaluation should answer.
In the future, testing AI systems will probably become as natural a part of the development process as unit tests, integration tests, or monitoring. Not because AI is “dangerous by definition”. Simply because a system whose output is not always deterministic requires a different way of measuring quality.
And the more AI moves from generating text to handling real processes, using data, RAG, and carrying out actions through agents, the more important it becomes not only to ask “can AI do this?”, but also: “can we prove that it does it well enough?”



