by Giuliana Parentelli
On Thursday, September 10, we had the chance to attend Testear.la 2026 | QA Conference in Buenos Aires, a day that brought together the software quality community to talk about testing, artificial intelligence, cloud, security, development, operations, culture, and, above all, how our own roles are changing.
If we had to sum up the entire event in one idea, it would probably be this:
AI can do things a lot faster. Including mistakes.
So no, apparently we still can’t leave agents working on their own while we go grab a coffee. Although there were moments during the day when it felt like our profession was at risk.
That’s exactly where the conversation at Testear.la went: we’re no longer just discussing whether we’re going to use AI. We’re starting to discuss how much autonomy to give it, how to control it, how to test it, and, above all, when we can trust what it does.
The day kicked off looking at the cloud, with the panel «Who Controls the Cloud?». The discussion centered on agent autonomy and left behind an idea that would repeat throughout the entire event: autonomy shouldn’t be handed over all at once, it should be earned progressively. First read, then recommend, then act under supervision, and only when the risk allows it, execute autonomously. For QA this also means a shift: with probabilistic systems there isn’t always a single correct result anymore, and the classic pass/fail starts to fall short.
Next, Nicolás Magni brought AI down to earth, talking about what happens when a nice demo has to survive in production. His proposal was concrete: take advantage of AI’s flexibility where it truly adds value, but surround it with deterministic rules, validations, observability, and traceability. Because an agent reaching the correct result isn’t enough — we also need to understand how it got there.
The panel «Leading the New Era of Quality Engineering» showed that AI is already part of QA’s everyday work: generating test cases, data, automations, evidence, and analysis. But the flip side also came up: now we have to learn to evaluate hallucinations, bias, security, and answers that can be different and still be valid. Agents tasked with auditing other agents are even starting to show up.
Yes, agents policing agents. All that’s missing is one that files the Jira ticket when something goes wrong.
One of the most interesting questions came from Rosana Suárez’s talk on code review: if AI writes the code and another AI reviews it, who validates the AI? Tools like CodeRabbit, Kiro, or PR Agent can speed up review enormously, but a green pipeline doesn’t guarantee a change is correct from an architecture, security, maintainability, or business standpoint. AI can help us review; judgment still can’t be fully outsourced.
And judgment was exactly one of the words that stood out in the HR panel. If everyone has access to tools that can generate code, tests, and documentation in seconds, the differentiator starts to be knowing what to ask for, what to question, what to discard, and how to apply what’s generated to the real business context.
There was also room for security. Sebastián Passaro presented «Does Your RAG Eat Whatever Comes Its Way?», showing that a RAG system can retrieve contaminated documents and end up giving false or dangerous answers. For QA this opens up a whole new playing field: prompt injection, malicious documents, output guards, and adversarial testing are starting to join our toolkit.
With Axel Labruna and AgentOps, another interesting idea came up: autonomy doesn’t mean authority. An agent can research, analyze, and generate proposals without having permission to directly modify a system. And the bigger the impact of an action, the stronger the controls should be. Even agents need onboarding: context, permissions, limits, and knowing when to raise a hand and ask for help.
The day also stepped outside the traditional software world to get into video game QA, where it’s not enough to check that something works — you also have to validate physics, balance, difficulty, narrative, matchmaking, and even something as subjective as game feel. A good demonstration that quality was never just about finding bugs.
And finally came the idea of the «augmented tester,» from Sebastián Layana: using AI as a copilot to understand new technologies, design better strategies, find scenarios we hadn’t considered, generate data, and improve our communication. But with an important warning: don’t turn into a «meat proxy» — someone who simply copies whatever AI returns and pastes it elsewhere without questioning it. The tool should enhance our judgment, not replace it.
And after a full day of talks, agents, automations, RAGs, generated code, and discussions about the future of QA, one question kept circling back…
After spending the whole day hearing about agents that code, agents that test, agents that review other agents, and agents that will probably soon be scheduling our daily standups, we came back with a few questions of our own.
Two in particular kept coming up in our conversations:
As we delegate more and more tasks to AI agents, how much of what they generate should we actually be reviewing?
And perhaps the bigger question:
What is the QA role going to become as all of this evolves?
The first part has a pretty clear answer after Testear.la 2026: you have to review it. And we’ll probably have to review more and more, though not necessarily line by line.
AI lets us generate code at a speed that makes it increasingly hard to review everything manually. And that’s where one of the challenges mentioned during the day comes in: the verification gap. We can produce software faster and faster, but that doesn’t mean we can validate at the same speed that what’s generated is correct, safe, and actually meets what we needed.
So maybe the question stops being just «did we review all the code the agent generated?» and starts being «what do we need to review, and what evidence do we need to trust what it generated?» Review starts moving up a level.
Maybe QA doesn’t just need to ask «is this code good?» but also «what evidence do I have that this system does what it’s supposed to do?» That’s where specifications, acceptance criteria, tests, architecture, observability, traceability, quality gates, security, and real production behavior come in.
In fact, another talk raised the idea that with Spec-Driven Development, code could become increasingly regenerable: humans define and maintain intent through specifications, and agents generate code, tests, and documentation. In that scenario, testing risks becoming the new bottleneck, because generating is cheap — verifying still costs.
And what does QA turn into? Maybe less «the person who finds bugs» and more the person who builds trust. A kind of Software Truster, as it was put during the event: someone who understands product and business, designs evaluation strategies, questions specifications, tests probabilistic systems, oversees agents, analyzes risk, demands evidence, and helps decide when a system is truly ready to move forward.
The tester of the future probably won’t have to compete against AI to see who writes more test cases per minute. That would be a pretty boring competition. And we’d clearly lose it anyway. Our edge is somewhere else: asking the questions AI doesn’t know it should be asking.
After a full day of talking about artificial intelligence, autonomous agents, automation, and code generated in seconds, there was one fairly paradoxical takeaway: as QA people… well, professional distrust has always been a bit of our superpower.
The more capable machines become at building software, the more important our human ability to distrust it wisely becomes.
With a 360° potential, our solutions matrix accompanies the lifecycle of any project, with skills and experience in Development, Design, Q&A, Devops, Operation & Deploy, and Architecture
We are here to help you!
You can leave us your query or recommendation through this form.
I accept the terms & conditions and I understand that my data will be hold securely in accordance with the privacy policy.