AI pentesting: finding is cheap now, proving is the job

Vulnerability disclosures doubled in 2026 and just 0.23% were ever exploited. Here is why a pentest report without a reproducible exploit is now triage work.

PentestBR TeamPublished 5 min read

In January 2026 the world received 5,045 vulnerabilities. In August it received 10,740. The count doubled in eight months, and the obvious reading is that software got twice as dangerous. It didn’t. Google Threat Intelligence published a study on September 30 about precisely this: of everything disclosed in 2026, 0.23% was ever observed being exploited. One in 431. What follows is the number that decides where your security budget goes, and why a pentest report that arrives without proof has quietly turned into triage work.

What Google Threat Intelligence actually measured

The GTIG report covers 20 months of disclosures and exploitation observations, January 2025 through August 2026. Three numbers carry the rest of this piece.

First, volume: 5,045 disclosures in January 2026, 10,477 in July, and a peak of 10,740 in August. Second, exploitation: 141 distinct vulnerabilities were both disclosed and exploited between January and August 2026, more than the 127 exploited across all of 2025. The monthly average went from 10.5 to 18.

The third number is the one that takes the list apart. Only 0.23% of vulnerabilities disclosed in 2026 were observed in active exploitation, roughly one in 431. GTIG is careful that raw volume misleads: automated CVE assignment in open-source ecosystems inflates the count. Vulnerabilities whose description merely mentions “Linux Kernel” generated about 5,000 CVEs between January and August 2026, with zero in-the-wild zero-days observed.

So a good chunk of what reached your desk as a “new problem” is not double the risk. It’s double the text.

Not every real bug is a bug that matters

Worth looking at the other side, because this study is not an argument against AI. Among the vulnerabilities GTIG could attribute to autonomous agents in 2026, the risk profile inverts: 58% are medium risk against 28% across the rest of the ecosystem, and low-risk findings drop from 69% to 39%. Exactly half of AI-discovered vulnerabilities lead to remote code execution, versus 26% overall. The report attributes that to how the work gets assigned: research programs point agents at privilege boundaries and multi-step code paths, which is exactly where the boring logic bugs live.

GTIG also warns that public data undercounts AI-discovered vulnerabilities, because there’s no standard attribution metadata, because cloud providers patch straight into production without requesting a CVE, and because many findings sit under embargo. In other words, the true AI-discovered total is larger than anything we can measure. That doesn’t change the conclusion. Whether the underlying number is measured, unmeasured, or estimated by a third party, the triage queue in front of you is long.

The case that ties both halves together is CVE-2026-1731: an unauthenticated OS command injection in BeyondTrust Privileged Remote Access, discovered autonomously by the Hacktron AI research agent. Four days after public disclosure, GTIG was already watching a threat cluster exploit it. Five more clusters within seven days, with privilege escalation, data exfiltration, and secondary payloads dropped on the way in. That was not a report left to rot in a folder. It was a door that opened fast, because somebody proved it opened.

Google turned off the report button

The same day, Google suspended product submissions to its open-source rewards program, the OSS VRP. The reason is in their own statement: the pause is due to a significant rise in automated submissions, “the vast majority of which are not valid.” Supply chain reports and already-outstanding reports are still accepted, and the company promises to rework the program by Q1 2027.

That program has paid out more than $81.6 million since 2010, including $17.1 million in 2025 to over 700 researchers. This is not a side project squeezed by budget cuts. The channel was closed because the cost of human time spent reading machine-generated submissions went past the return. And it isn’t the first: in January the curl maintainer shut down his own program on HackerOne for the same reason, and in September Intel removed the financial rewards from its Intigriti program without explanation.

A very large technology company arrived at the conclusion I’ve been arguing for years: generating a report is cheap, verifying a report is expensive, and the market only pays for the second half.

What this has to do with your next pentest

This is where the news turns into a buying decision, and where I will argue against the standard advice.

If you adopt a scanning tool because it “found 400 vulnerabilities,” you didn’t buy security. You bought a queue. By GTIG’s own numbers, every 431 items in that queue contains one somebody will actually exploit, and the other 430 are triage cost, report noise, and, if they pass through unfiltered, noise for you and for your auditor.

Two things I’d say to a CTO in that spot. First: the test scope has to deliver proof, not a headline. A serious pentest finding comes with a reproducible request, an observed response, and the effect demonstrated in your environment. That’s what separates an audit-ready report from a forum thread.

Second: the count of findings is not a security metric. The security metric is what the attacker managed afterward. FortiMail, Cisco, and NetScaler all put on the same show this week: patches published, appliances rebooting themselves because somebody was already inside, and admins reporting they patched and still got hit. Updating closes the entry point of one particular path. It does not close a session someone already opened, revoke a credential that already left, or tell you whether a second door exists.

If I were that CTO, the first thing I would do tomorrow is not swap scanners. It would be to go down the queue and demand a reproducible proof for every open finding, in a test environment. Whatever survives that is your real backlog. The rest can wait for next quarter.

What to do now

  • Ask for evidence, not counts. Every finding needs a request, a response, and an observed impact in your environment.
  • Treat a scanner list as a triage queue with an owner and a deadline, not as a fix backlog.
  • Measure the test by what the attacker did afterward, not by how many items shipped.
  • Put privilege boundaries in scope, where AI agents are finding double the remote code execution rate and where manual testing usually skims past.
  • Re-test after the patch. Updating does not clean up an earlier intrusion.

If your web and API surface has not been tested end to end this year, that is where the PentestBR coverage scope comes in: manual testing with approval, reproducible exploits, and a report your auditor accepts. You are entitled to argue with the 0.23%, and I am not claiming your company has 431 vulnerabilities waiting to happen. What I am saying is this: when someone hands you a very long list, ask how many items on it were proven in your environment. The answer is usually much smaller than the list.

Keep reading

Back to blog