approx. 10 min. reading time
Software Quality Days 2026: When AI Tests, Humans Must Decide
Written by Florian Fieber /
June 2026

Table of Contents:
AI generates code, AI runs tests, all tests pass – and yet responsibility for what goes into production remains with the human.
At the panel “What, still testing ourselves?” at the Software Quality Days 2026 in Vienna, three experts joined me to discuss what AI genuinely changes in software testing – and what it does not. This article summarises the key positions: on hype and reality, on accountability, on the junior-senior problem, and on how we preserve human judgement in a world where it becomes increasingly easy to do without it.
At the panel “What, still testing ourselves? We now have AI for development and testing!?” at the Software Quality Days 2026 in Vienna, three experts joined me to discuss what AI genuinely changes in software testing – and what it does not. This article summarises the key arguments and positions.
An early-morning question to the audience, just after eight o’clock: “The early bird catches the…” – the room completes the sentence before I finish speaking. That is exactly what large language models do. They complete patterns – sentences, code, test cases, entire software systems. The code turns green, the tests pass, the software goes to production. The question discussed for the better part of an hour: what does that mean for the people who used to be responsible for all of this?
The Panel
The discussion was a joint initiative of the German Testing Board, the Austrian Testing Board, and the Swiss Testing Board. I had the privilege of moderating. Three panelists with very different perspectives on the topic joined me on stage:
Dr. Elmar Jürgens, founder of CQSE and a leading researcher in software quality analysis – with the rare advantage of testing research insights daily against the reality of the companies that use his products.
Mateusz Gren, Head of Software Development at Atos Austria, who works with AI-assisted development not as a concept but hands-on, every day, in projects with hundreds of thousands of lines of code.
Lucy Dinu, founder and CEO of KHAIO, with a strategic background spanning NATO, Daimler, and other organisations – bringing a perspective on AI that starts not with the technology but with the question of what organisations actually want to achieve with it.

Do We Still Need Human Testers?
Lucy Dinu’s answer was direct: yes. Human judgement and critical thinking cannot be automated. Mateusz Gren agreed: humans remain in the lead, but how software development works is changing fundamentally.
The most nuanced answer came from Elmar Jürgens:
“I think in a testing conference it’s very tempting to now say that we as testers will always be needed. And I think that’s simply not true, at least for a certain definition of tester.”
His argument: the word “computer” referred to a human being performing mathematical calculations for over three hundred years. What we mean by “tester” today is already different from twenty years ago – and in twenty years it will have shifted again. Those who clicked through scripts manually have largely been replaced in many areas. What remains is what an LLM cannot autocomplete: critical thinking, contextual understanding, asking uncomfortable questions.
This distinction matters more than the reassuring blanket answer. The role is changing – that is a design challenge, not a threat. But only if we actively engage with it.
Hype vs. Reality: What Actually Works
As this year’s Program Chair of EuroSTAR in Oslo, Elmar Jürgens has reviewed all submissions for the conference programme. Of 550 submissions, roughly 300 were about AI in testing. The proportion of those who brought genuine hands-on experience – concrete results rather than promises – was significantly smaller. His principle: he only believes claims about what AI can do if he knows someone who has personally done or observed it.
Mateusz Gren brought figures from practice: Atos measured the productivity impact of AI in 2025 across five roles – project managers, testers, software developers, requirements engineers, and architects. The results were highly context-dependent. Sometimes working without AI assistance is simply faster. And measurement itself quickly hit structural limits: much like the introduction of MS Office, the contribution cannot be cleanly isolated – it becomes part of how work functions.
A metaphor I had read that very morning captures the problem well: imagine a motorway full of cars stuck in a traffic jam. Now replace every car with a Ferrari. The jam remains. The bottlenecks are not in the vehicle – they are in architecture, requirements, coordination, and quality assurance. Generating more code faster does not solve that. It creates more that needs to be understood, tested, and maintained.
This is exactly where the paper by Margaret-Anne Storey becomes relevant. In “From Technical Debt to Cognitive and Intent Debt,” she proposes a Triple Debt Model. Alongside the familiar technical debt – shortcomings in the code itself – she identifies two new forms of debt that arise from AI-assisted development. Cognitive Debt describes the erosion of shared system understanding within a team when code is produced faster than people can comprehend it. Intent Debt describes the loss of explicitly documented rationale – the knowledge of why a decision was made a certain way, which is then missing for both humans and AI agents working with the system later.
The model explains why productivity promises and practical experience diverge so widely: the gains at the surface are real. The debt that accumulates is harder to see – and will be paid eventually.
Where Is the ROI? Three Perspectives
Lucy Dinu describes an approach in her consulting practice that begins with strategy rather than technology:
“If the strategy is just ‘give me AI so I can increase my ROI’, that’s not really a strategy.”
Her example is concrete: if a marketing colleague saves five hours per week through AI – what happens to those five hours? Without strategic grounding, the gain evaporates. And sometimes the boldest question to ask is: do we actually need AI here at all?
Elmar Jürgens differentiated by error consequence: the hype concentrates on areas where bugs carry limited cost – startups with short runways, internal tools, throwaway prototypes. In heavily regulated domains such as automotive or financial infrastructure, adoption is significantly more cautious – and for good reason.
His own observation at CQSE: since introducing Claude Code in January, the team has produced noticeably more activity in non-critical internal systems. In the core product? No measurable change in code volume – though qualitatively the feeling of moving faster. That is not a data basis that justifies 10x promises.
Accountability: Who Is Responsible?
When AI writes the code, AI runs the tests, all tests pass – and then a serious defect appears in production: who is responsible?
Lucy Dinu used the Michelin restaurant metaphor: the head chef is responsible for everything that leaves the kitchen, regardless of who plated the dish. With the EU AI Act, software is increasingly treated as a product subject to liability. AI tools are instruments; instruments do not bear responsibility.
Mateusz Gren made it concrete at the process level:
“If you do automated code reviews, it’s the architect who put that into place – he’s responsible. If it is a developer using Claude for writing his unit tests, and they are nonsense, it’s the developer who committed this unit test.”
Elmar Jürgens added an engineering perspective: the question “who is to blame?” typically incentivises the wrong behaviour – developers gravitate toward trivial tasks to avoid being at fault. More effective is anchoring accountability at the process level. CQSE has always done full code peer reviews, ensuring that at least two people are accountable for every line. That creates quality without blame. Anyone deploying AI tools takes responsibility for their output. In a regulated world, that is the only legally and ethically defensible position.
The Junior-Senior Problem
The logic many organisations are currently following: we need fewer people, we only need seniors who can orchestrate AI. Juniors, who learn slowly and require significant mentoring, are no longer cost-effective.
I introduced a thesis into the discussion: “The new entry level is a senior.” The idea: through AI, early-career professionals can operate at a higher level from the start, because foundational work is automated.
Elmar Jürgens pushed back with a simple argument: CQSE has hired almost exclusively juniors throughout its 17-year history – and sees no reason to change that. The capacity to learn, curiosity, the willingness to put in effort – these are the qualities that count over the long term. Whether someone already knows requirements engineering is learnable. Whether someone wants to learn is not.
Lucy Dinu added: the generation that has grown up with AI tools sees things differently – and that different perspective is valuable. Not despite their inexperience, but sometimes precisely because of it.
Mateusz Gren named the challenge from a business perspective directly: clear development paths for juniors are missing. The role is changing so quickly that organisations themselves do not yet know what career trajectories will look like in ten years. That is not an excuse – it is a collective task for the entire industry.
If we stop investing in juniors now, we are sawing off the branch on which the next generation of seniors will sit. AI generates code today. But whoever decides in ten years whether that code is good must have learned that at some point – through real experience, not through prompts.
What to Do? Three Closing Messages
Lucy Dinu: stay curious, actively question AI, do not trust it blindly. And do not neglect the human part. The capacity for independent thought atrophies if it is not exercised.
Mateusz Gren: if you are not experimenting with AI tools for at least thirty minutes every day, you are falling behind. The role is changing – those who do not shape that change will be shaped by it.
Elmar Jürgens said what I consider the most important thing from the entire panel:
“Don’t stop doing as much quality assurance as you do now because it feels like we’re just moving faster – so more things will go wrong, so we can’t do less of the things we used to do to keep things from going wrong.”
More speed does not mean less need for quality. It means the opposite.
Conclusion: The Wrong Question, The Right Task
The panel began by asking whether we still need human testers. The more honest question is: how do we protect human judgement in a world where it becomes increasingly easy to do without it?
AI can write code, generate test cases, and detect anomalies. It cannot decide what a piece of software should be permitted to do. It cannot bear ethical responsibility when an algorithm disadvantages people. It cannot assess whether a system truly serves the purpose for which it was built.
That is our task – the testing community’s, quality engineers’, and all those who understand what good software means. Not only what working code means.
For the three DACH-Boards ATB, GTB and STB, this is not an abstract debate. It is the question of which competencies we embed in certifications, which topics we prioritise in curricula, and which discussions we continue – together with our partner boards in Austria and Switzerland – at conferences like the SWQD. The conversation in Vienna was a good start. The work on it has only just begun.
References
| [1] | Margaret-Anne Storey, “From Technical Debt to Cognitive and Intent Debt: Rethinking Software Health in the Age of AI”, arXiv:2603.22106, March 2026. https://doi.org/10.48550/arXiv.2603.22106 |



