1 min read
AI on Trial – Insights from EuroSTAR 2025
In May 2025, the EuroSTAR conference once again brought the European software testing community together - this time in Edinburgh under the motto "AI...

Practical. Proven to success. Tailor-made. Learn more about our case studies.
10 min read
Sabri Deniz Martin
:
Friday, 4.9.2026
Financial institutions are increasingly integrating AI into business-critical processes. This not only creates new opportunities but also raises the bar for data quality, governance, traceability, and software testing. Here are our key takeaways from Frankfurt.
The Handelsblatt Banking Summit and the AI in Banking conference took place simultaneously on September 2 and 3, 2026, at Kap Europa in Frankfurt am Main. With a single ticket, attendees could switch between the two programs; individual masterclasses and thematic sessions were held across both conferences. While the Banking Summit focused on the strategic, economic, and regulatory changes in the industry, “AI in Banking” concentrated on specific AI applications, scaling, governance, and the organizational prerequisites for productive deployment. The official programs reflected this close integration: the range of topics spanned from AI-supported product development to KYC and credit processes, as well as the data economy and agent-based systems. (AI in Banking Program, Banking Summit Program)
Enrico Ausborn, Head of the Financial Services Business Unit, and Sabri Deniz Martin, CMO, represented TestSolutions at the event. During the presentations, panel discussions, and conversations, they focused in particular on the requirements that the productive use of AI places on processes, data, and quality assurance.
Our strongest impression after two days: The industry is no longer merely debating whether to use AI. What matters now is how financial institutions can turn promising pilot projects into robust, cost-effective, and audit-proof systems. Banks took center stage at both events, but many of the challenges discussed are equally relevant to other financial service providers. As a result, AI in the financial sector is increasingly becoming a matter of quality.

Enrico Ausborn, Head of Business Unit Financial Services, and Sabri Deniz Martin, CMO.
During the panel discussion “From Use Cases to Scaling: How Banks Can Sustainably Integrate AI into Their Organizations,” it became clear that successful AI implementation cannot take place in isolation within IT. Strategy, business units, IT, data teams, governance, and stakeholder involvement must all work in unison.
Florian Mielke, Head of the AI Office at DZ BANK, emphasized the importance of central organizational anchoring and taking stock of the AI systems in use—including their cloud environments and costs. It is equally important to empower employees and take employee participation seriously. Dr. Gerhard Svolba of the SAS Institute pointed out that many of the best use cases are likely to come from the business units. Philipp Schwaab, Chief AI Officer of the Helaba Group, in turn highlighted data quality as an essential foundation.
A pattern is thus emerging: concrete use cases arise from the bottom up, while the mandate, platform, guidelines, and prioritization must be established from the top down.
A Center of Competence or AI Office can manage a use-case funnel, establish standards, and support the business units in implementation. However, without the backing of top management and without clear responsibilities, many projects remain stuck in the pilot phase.
Sabri Deniz Martin, CMO at TestSolutions, summarizes the necessary diligence as follows:
“Successful AI projects begin with selecting the right use cases and defining clear metrics for quality, risk, costs, and benefits. This requires coordination and project planning. Companies must then create the conditions for implementation and give the systems time to learn. AI agents can certainly be viewed as new employees in this context: They need a clear scope of responsibilities, good information, feedback, and oversight. Above all, they need a reliable data foundation. Anyone who lacks these fundamentals — or is unable or unwilling to establish them — should critically reevaluate the productive use of AI at this time.”
The question has been raised repeatedly: How can the economic benefits of AI be assessed in complex end-to-end processes? The potential is particularly tangible where a process is executed frequently and scales well — such as in customer service, Know-Your-Customer checks, credit processes, or the software development lifecycle.
Using the credit business as an example, it became clear that AI does not come into play only during the credit decision itself. Potential exists throughout the entire lifecycle: in the selection of suitable borrowers, in faster and more consistent underwriting decisions, in the early detection of deterioration, in the prioritization of significant exceptions, in the reduction of recurring manual checks, and in feeding actual credit outcomes back into a learning loop. The expected value is distributed accordingly across growth, risk reduction, efficiency, and risk-adjusted capital allocation.
The key unit of analysis here is not just the model, but the entire workflow. For each process step, it must be clarified what contribution AI can actually make, which parts should continue to be automated deterministically, and which patterns can be applied to other processes. Added to this are costs for infrastructure, model usage, integration, operation, monitoring, and data preparation.
Martin Flechl of Dynatrace drew attention to another point: A model that is nominally more expensive can be more cost-effective if it achieves better results with fewer queries. Model prices or token costs alone, therefore, are not very informative. Costs, response quality, latency, error rates, and business value must be measured together.
A presentation on the use of AI in credit risk and lending illustrated this concept as a continuous AI FinOps cycle. This includes transparency regarding expenses, usage, quality, and results per workflow; the optimization of unnecessary calls and oversized contexts; clear budgets and responsibilities; as well as forecasts and benchmarks for costs per unit of work. Among the levers mentioned were model routing, a redesign of workflows, improved context and data engineering, and domain-expert-curated “golden answer” evaluation sets. For example, small models can handle extraction and validation, while more powerful models are used only where synthesis is actually required.
Consequently, the ROI of an AI system is not a one-time business case calculation. It is an ongoing operational metric that must combine technical and domain-specific quality metrics.
Data quality was one of the recurring themes throughout both days of the conference. When information is scattered across many systems, documents, and formats, significant effort is required even before the actual AI deployment begins. Data must be located, made accessible, structured, evaluated, and placed in a context that the system can utilize.
This has a direct economic consequence: Poor or fragmented data not only increases the risk of errors but also the implementation and operating costs of an AI use case. Additional retrieval steps, repeated model calls, manual rework, and more complex checks drive up the effort required.
For quality assurance, this means that an AI system cannot be meaningfully tested without simultaneously verifying the quality of its input data, knowledge sources, and interfaces.
Data quality thus becomes an integral part of the test strategy, rather than an upstream task that can be considered complete at go-live.
Several presentations focused on so-called digital workers or AI agents. Such an agent combines a language model with tools, interfaces, retrieval, context, and clear guidelines. It is not only meant to generate a response but also to deliver results within a defined scope of tasks, utilize existing systems, escalate to humans in cases of uncertainty, and document its work in an audit-proof manner.
An IBM masterclass contrasted the classic, sequential workflow with an agent-based control loop. Instead of a rigid sequence of input, rule checking, processing, approval, and output, agents are expected to understand and plan tasks, execute steps, review results, create decision templates, and adapt their approach based on feedback. Humans remain “above the loop”: they oversee and approve the process and intervene in the event of escalations. For an illustrative KYC process, IBM contrasted a traditional processing time of several hours with significantly shorter automated preliminary work followed by human review. Such figures should be viewed as examples provided by the vendor and not as a universally applicable performance promise.
A step-by-step approach was outlined for implementation: First, existing processes are analyzed and optimized. They are then broken down into autonomous, automated, and manual individual steps. Building on this, the required skills, MCP interfaces, system integrations, guardrails, and “human-in-the-loop” points can be determined. The key prerequisite for this is robust process documentation.
However, this is precisely where one of the most challenging quality issues lies: An agent cannot reliably detect or signal uncertainty in every case. Escalation to a human is therefore only as good as the criteria by which the system identifies uncertainty, contradictions, or a lack of evidence.
A feedback and learning loop must also be controlled: Which corrections are allowed to be fed back into the system, who validates them, and how do we prevent erroneous decisions from being reinforced? Guardrails alone do not solve this problem.
Yiping Li, Global Chief Operating Officer of Deutsche Bank Private Bank, provided a particularly concrete example in the case study “AI-Driven KYC Processes.” In private banking, complex asset backgrounds must be traced, numerous documents evaluated, and regulatory requirements met. According to the presentation, approximately 48 percent of the processing effort is attributable to client due diligence and the verification of the source of funds.
The end-to-end analysis presented ranged from preparation and data collection, through the creation of the client profile and risk rating, to due diligence, the KYC memo, and approval. Right from the start, a guided digital data collection process can standardize data and pre-fill client profiles with existing information. In later steps, AI can review documents, flag anomalies, prioritize screening hits, reduce false positives, and generate KYC memos from validated data and approved templates. The presentation focused on source-of-wealth verification within client due diligence.
The multi-agent approach presented distributes tasks among specialized components: one agent structures unstructured data, another evaluates documents such as older pay stubs, and other components cross-check client information against registries, publicly available information, and internal data. The result is intended to include, among other things, an evidence-based “Wealth Accumulation Journey” with confidence scores for individual verification steps. Humans remain continuously involved in the process.
This use case in particular demonstrates why traditional test logic alone is insufficient. A technical test can pass even though a business-related conclusion is implausible. For example, if a system infers supposedly typical language proficiency from a job profile, the output is syntactically correct and may even be statistically plausible—but it is still incorrect for that specific person.
For us, one of the key takeaways from the presentation is this: AI projects require a great deal more testing. In traditional IT projects, IT performs certain tests and business analysts perform others. But with AI projects, the tests may pass — and you still need to test much more because the output is non-deterministic. You need more test cases and greater variability.
The result is closer collaboration between IT testing, business units, compliance, second-line oversight, audit, and, where applicable, regulatory bodies. In addition to functions and interfaces, factors such as business plausibility, evidence-based decision-making, robustness, variance, biases, traceability, and the quality of human-in-the-loop decisions must be evaluated.
Equally pointed was the recommendation not to make AI an end in itself:
“Wherever you can use deterministic logic, do not use AI.”
Traditional automation is more cost-effective, more reproducible, and easier to control in many process steps. The most viable architectures are therefore likely to be hybrid: deterministic logic where rules suffice, and AI where unstructured information, complex patterns, or linguistic assessments must be processed.
A layered model developed by the Business Engineering Institute St. Gallen revealed that financial institutions do not build their AI on a single technology. The chain extends from energy, data centers, and chips, through cloud infrastructure, standards, and protocols, to data, models, orchestration, and the actual applications.
The degree of design freedom is unevenly distributed: when it comes to their own use cases, applications, and data, institutions have relatively significant autonomy; however, when it comes to orchestration platforms, base models, cloud offerings, chips, and energy supply, there are greater dependencies on external providers.
For quality assurance, this means that an AI system should not be tested against just a single model. Model changes , switching cloud providers or service providers, new versions of orchestration frameworks, and changes to external APIs can all influence the behavior of the entire system. Portability, interchangeability, failure scenarios, and regressions across multiple layers must therefore be included in the testing strategy. This also applies to open protocols and agent interfaces such as MCP or A2A.
The EU AI Act, DORA, and GDPR formed the regulatory backdrop for many discussions. These discussions focused not only on guidelines and governing bodies but also on how requirements can be monitored during ongoing operations.
The “AI Control Tower” is described as a possible operational model that, among other things, oversees approved models and products, usage rules, guardrails, monitoring, costs, and risks. For financial institutions, it is crucial not to treat governance as a static collection of documents. Requirements must be translated into verifiable controls, technical thresholds, approvals, logs, and escalation procedures.
This includes monitoring model and system behavior. Which sources were used? Which tools did an agent access? What permissions did the agent have? When was a human involved? Was the result modified or discarded?
Auditability can only be achieved if such questions can be answered based on reliable evidence.
Dr. Andrea Hammermann from the German Economic Institute examined the impact of AI on work and organizations. Knowledge-intensive services are already making extensive use of AI, although younger employees, men, and managers have been working with it more frequently so far. Many employees perceive an increase in work performance; those just starting their careers, in particular, experience a steep learning curve.
At the same time, the availability of a tool does not automatically lead to higher productivity. If processes remain unchanged, AI is often merely superimposed onto existing work steps. “Shadow AI” arises particularly where employees are not provided with suitable or sufficient official tools.
In our view, this raises a key question for change programs, one that is also relevant for quality assurance and risk management: Which processes and decisions do we want to organize differently in the future, and what level of risk are we willing to accept?
This determines specific decisions regarding time requirements, quality standards, error tolerance, human oversight, and responsibilities. Involving employees and ensuring their participation early on is therefore not merely a supporting measure, but a key factor for success.
Andreas Schranzhofer, CTO of Scalable Capital, demonstrated how banking functions could be accessed in the future not only via websites and apps but also through AI assistants, command-line interfaces, and the Model Context Protocol. Depending on its role and use case, an agent receives information, skills, and tiered read and write permissions, and accesses existing functionalities via APIs.
This changes more than just the user interface. For example, when an agent analyzes multiple positions and prepares or executes stop-loss orders, identity verification, permissions, transaction limits, confirmations, and revocation options must function consistently throughout the entire chain.
Traditional API and end-to-end tests remain indispensable but are supplemented by new attack vectors and error patterns: faulty tool selection, prompt injection, privilege escalation, incomplete context, or unexpected action chains.
The two conferences highlighted a financial sector that is shifting from experimentation to operationalization. The most promising use cases lie in KYC, credit processes, customer service, analytics, software development, and operational management. At the same time, as systems become more autonomous, so does the responsibility for their quality.
For financial institutions, it is not enough to evaluate a model and then integrate it into a process. The entire socio-technical system must be tested: data and retrieval, models, prompts, traditional software, APIs, roles and permissions, process logic, human-in-the-loop points, monitoring, and audit trails.
The key question, therefore, is not whether an AI delivers a convincing answer in a demo scenario. It is: Does the overall system consistently deliver sufficiently good, secure, cost-effective, and traceable results under realistic conditions—and can this be proven?
Enrico Ausborn, Head of the Financial Services Business Unit at TestSolutions, sees this primarily as an opportunity for the industry:
“Many financial institutions have clearly moved beyond the phase of purely experimental AI. Now the focus is on rolling out viable applications on a broad scale—with clear business value and quality assurance that integrates functional, technical, and regulatory requirements. If institutions take an end-to-end view of their processes and prioritize quality from the outset, AI can not only boost efficiency but also make services faster, more consistent, and more customer-friendly.”
This is precisely where the major topics of the Banking Summit intersect with the core mission of professional quality assurance.
Do you have questions about AI?
Get in touch.
1 min read
In May 2025, the EuroSTAR conference once again brought the European software testing community together - this time in Edinburgh under the motto "AI...
1 min read
EuroSTAR is Europe’s largest software testing conference, with over a thousand attendees. This year, TestSolutions was represented by two speakers:...
1 min read
The General Data Protection Regulation (GDPR) is having a significant impact on how companies test software. While GDPR compliance may seem...