The inconvenient truth about AI pentesting: someone has to check all the work

Wait 5 sec.

AI pentesting can flood teams with findings they cannot validate. The real challenge is managing “validation debt” as discovery scales.AI pentesting has a ‘Sorcerer’s Apprentice’ problem. Enchant a broom to fetch water, and it will fetch water, relentlessly, long after the workshop has flooded.The industry is busy measuring how fast AI finds vulnerabilities (we’re looking at you, Anthropic). Far fewer people are costing out who checks all that work. That gap has a name: validation debt, and it’s the backlog of unverified findings that rolls in when discovery scales and verification doesn’t.In a recent survey, 158 practitioners were asked whether their teams could triage more than 500 AI-generated vulnerability candidates from a single engagement. Only 20.3% said they had a workflow in place to handle it. Another 38.6% said that volume would strain the team, and 29.7% said it was unmanageable.That means most teams are buying discovery capacity they cannot process.Finding more vulnerabilities is useful only if the team can work out which are real, which are important, and what to do about them.What happens when AI finds more than your team can handle?Security teams have spent years trying to test more applications, more often. AI finally makes that possible. The problem is that discovery scales much faster than the work that follows it.Someone still has to reproduce the finding, establish whether it is exploitable, and give engineering enough evidence to act on it.At low volumes, this works. With hundreds or thousands of findings, it’s overwhelming, if not impossible.One respondent spent two days validating 300 findings from an AI tool. 250 of them were duplicates, non-exploitable issues, or references to vulnerabilities that did not exist.Two days of manual validation is not the efficiency gain the team bought the tool for.Is AI saving time, or moving the work elsewhere?For some teams, perhaps. But much of the manual work is kicking the can down the road.81.7% of practitioners using AI tools discovered findings that needed significant manual validation ‘at least sometimes’.If automation cuts 10 hours from vulnerability discovery but adds 15 hours of validation and triage, where’s the ROI? All you’ve gained is validation debt.Discovery is relatively easy to scale. AI can probe potential weaknesses and generate candidate findings in a fraction of the time it would take a human tester.Every additional finding, however, gives the team something else to check.Validation is not scaling at the same rate, because a plausible finding still needs evidence.Some vulnerabilities will turn out not to be exploitable, while others will duplicate the same underlying problem. Severity alone also tells you very little if the affected asset does not endanger the business.How much does a bad finding cost?Every low-confidence finding takes time to discard.False positives are usually discussed as an issue of accuracy. For security leaders making investment decisions, the hours they consume are just as important.At five minutes each, validating 1,000 findings takes over 80 hours. And five minutes is optimistic, it depends on whether analysts need to reproduce an issue, understand its context, establish whether it is exploitable, or all the above.It gets worse when the output seems authoritative.AI-generated findings look convincing thanks to polished descriptions, severity ratings, attack narratives, and remediation advice. None of that is proof.Someone still has to check that the finding is real before asking developers to fix it. The tool price is easy to see. The hours spent checking its work are not.Why do some teams manage the AI flood better?Testing maturity makes a difference.Among teams conducting fewer than five tests per month, only 4% reported having a formal workflow for handling high volumes of AI-generated findings, while 55% said that such volume would be unmanageable.Teams that test more cope better with the volume. They have better workflows and are less likely to get buried in findings.The sample gets smaller at the highest testing frequencies, so treat this as a trend, albeit a telling one.Test often enough, and you have to get good at triage – someone needs to own it. Findings need to meet a standard, and junk must stay out of the queue.Infrequent testing can hide weak processes, which AI-generated volume exposes very quickly.Can your team handle what your AI tools produce?Start with the team’s capacity to deal with the output.Security leaders weighing up AI pentesting tools need to know how much work their current testing creates, and what would happen if that workload suddenly exploded.A few questions can help find the gaps:How many findings can your team realistically validate every week?What evidence must accompany a finding before engineering will accept it?How much analyst time goes into validating findings that lead nowhere?What happens if testing output increases fivefold but remediation capacity does not?Put those costs into the business case before declaring any efficiency gains (or parroting the ones from your preferred vendor).Generating ten times as many findings does not increase security capacity tenfold if the team cannot process them. It creates a bigger queue.Do you need more people or a better process?Both, but not in equal measure. Better processes can help teams absorb more findings, but they cannot make expert validation free.Hiring more analysts is not the only answer. Cutting avoidable work is a good start.Define the threshold that determines what qualifies as ‘validated’. Deduplicate ruthlessly so that one weakness does not become many tickets for different technical reasons. Prioritize based on risk; don’t treat every technically valid problem as the same.Testing frequency also counts. Teams that test regularly get more opportunities to improve how findings move from intake through validation, remediation, and retesting.If every AI-generated finding still needs substantial expert investigation, higher discovery rates will inevitably mean more labor. That cost belongs in the economics of automation.Stop measuring how much AI findsWe’ve always judged security tools partly on how much they can find. AI can push that number way up, but finding more doesn’t tell you whether a security team is getting more done.A better measure is what happens to those findings.How many are validated quickly? How many lead to remediation? How much human effort does it take to get there?Those are the questions that tell security leaders if AI is helping them gain efficiency.The cost of AI pentesting doesn’t stop at the license fee. It includes the people and processes needed to validate what the technology finds. Ignore that cost, and automation will create more work than it removes.About the author: Kirsten Doyle has been in the technology journalism and editing space for nearly 24 years, during which time she has developed a great love for all aspects of technology, as well as words themselves. Her experience spans B2B tech, with a lot of focus on cybersecurity, cloud, enterprise, digital transformation, and data centre. Her specialties are in news, thought leadership, features, white papers, and PR writing, and she is an experienced editor for both print and online publications. She is also a regular writer at Bora.Follow me on Twitter: @securityaffairs and Facebook and MastodonPierluigi Paganini(SecurityAffairs – hacking, newsletter)