5 min read
AI Evals Explained: Evaluating LLM Outputs and the challenges involved
If you've been following news on technical developments in AI, you'd have probably seen the term 'evals' suddenly showing up everywhere. In this...

Practical. Proven to success. Tailor-made. Learn more about our case studies.
On the evening of August 25, 2026, TestSolutions hosted an event at the Dorian Gray Lounge in the Frankfurt Airport Club. The venue was chosen deliberately: The Frankfurt Airport Center is located right where the aviation industry operates every day, and that’s exactly where our guests came from—specialists and executives from the airline sector, many of whom we’ve worked with for years on quality assurance and software testing.
Two presentations, two perspectives, one common thread: You should only rely on AI once you’ve verified exactly what it is you’re relying on.
Prof. Dr. Marco Barenkamp opened the evening with an uncomfortable question.
Prof. Dr. Marco Barenkamp, LL.M., a member of the TestSolutions Advisory Board, kicked off the event. As an honorary professor of artificial intelligence and digitalization at Osnabrück University of Applied Sciences and vice chair of the Federal Commission on Artificial Intelligence and Value Creation 4.0 of the German Economic Council, he combines academic depth with business practice.
He began with a phrase that many in the room recognized from their own daily work: “That’s what our AI does.” What initially sounds like efficiency turns out, upon closer inspection, to be a gap in accountability. When a result can no longer be attributed to anyone, ultimately no one checks it anymore.
Barenkamp then moved on to the central question of his talk: Are people losing their ability to think critically? He referred to a study by Microsoft Research and Carnegie Mellon University on the influence of generative AI on the critical thinking of knowledge workers. The effect described there is unsettling. The greater the trust in an AI system, the less cognitive effort people invest in verifying the result. Those who trust the machine check it less. Those who check it less realize later that something is wrong.
For organizations, this raises not a question of technology, but one of design: Where does professional judgment remain anchored in the process when systems formulate answers with ever-greater plausibility? The provocative phrasing of the talk’s title, “AI or Just Average Intelligence,” aimed precisely at this point. A model that reproduces the statistical average is an excellent assistant but a poor final decision-maker.
In the second part, the discussion got down to specifics. Florian Fieber, Chief Process Officer at TestSolutions, explored the question of how to assess whether an AI assistant is truly trustworthy. His example: a chat assistant in an airline’s service center.
Five error scenarios from a real-world test set for an airline assistant.
From the test set, he highlighted five typical error scenarios that occur regularly in operation:
Each of these errors can result in direct costs, operational rework, reputational damage, or legal liability. And each one is recognizable if you look for it. That’s the real point: All five answers sound plausible. Only those who know the underlying rule and look it up can identify them.
The case of Moffatt v. Air Canada as an example of structural risk.
Fieber demonstrated just how quickly plausible misinformation can turn into a legal problem using the case of Moffatt v. Air Canada, decided in February 2024 by the Civil Resolution Tribunal in British Columbia. Following the death of his grandmother, a passenger had inquired with the airline’s chatbot about bereavement fares and was told that the reduced fare could also be applied for retroactively. He booked at the full price. The airline subsequently rejected the request because retroactive applications were prohibited under its policy. The tribunal ruled that the company is responsible for all information on its website, regardless of whether it comes from a static page or a chatbot.
The direct damages were in the three-digit range and were economically insignificant. The other two aspects are interesting: The reputational damage resulted less from the decision itself than from the defense strategy used in the proceedings. And the structural risk remains: Every statement made by a chatbot is a statement by the company, and the risk scales with the volume of responses.
So how do you evaluate a system that generates thousands of responses every day? Fieber demonstrated the approach we take in our projects: subject matter experts evaluate a sample of responses based on clearly defined criteria. This evaluation is then scaled to large volumes of responses using LLM Judges. This transforms sporadic gut feelings into a repeatable process with reliable metrics.
The quality of the benchmark itself is crucial here. A benchmark that doesn’t accurately reflect the domain rules produces good metrics for a poor system—or vice versa. The key takeaway of the evening: Sometimes the problem isn’t the AI, but the benchmark we use to evaluate it. This isn’t an AI issue, but a testing issue—and thus precisely the field in which we’ve been working for nearly two decades.
Florian Fieber is responsible for test process management, product management, and the TestSolutions Academy at TestSolutions. He is Chairman of the German Testing Board (GTB), where he leads the GenAI working group and heads the ISTQB working group “Testing in Particular Domains.”
Afterward, the evening moved on to what it was intended for: professional discussion over drinks in small groups. The two presentations provided a solid foundation for this, as they addressed the same question from two different angles: Barenkamp from the perspective of the organization and the people within it, and Fieber from the perspective of verification.
We would like to thank all participants for the open discussions and both speakers for their presentations, which complemented each other perfectly.
Evaluating AI assistance systems in direct customer contact is one of our current priorities in the Aviation Business Unit. If you’d like to be informed about upcoming events or are interested in a professional exchange on quality assurance for AI systems, please feel free to contact us.
Not sure how well your AI assistant actually performs?
We’ll work with you to assess your quality risks before discussing the scope of the project.
5 min read
If you've been following news on technical developments in AI, you'd have probably seen the term 'evals' suddenly showing up everywhere. In this...
1 min read
The airline industry is facing an epochal change. With the IATA-driven transition to airline retailing systems such as NDC (New Distribution...
1 min read
In March 2026, TestSolutions GmbH took part in the Aviation Festival Asia in Singapore, one of the most important international conferences for...