AI pentest: the bug found in a day, exploited the next

CVE-2026-61500 in Rejetto HFS was found by an AI agent and exploited within 24 hours. The lesson is in the boring part of the code, not in the magic.

PentestBR TeamPublished 5 min read

The picture is tempting: an AI model finds a bug in an open source project, turns a math problem into a working exploit, and less than 24 hours later somebody is using that door in the wild. I get the temptation, it is a good story. It also hides the part your CTO actually cares about, and that part is far less glamorous: the flaw was not an AI flaw. It was a bad random number generator behind a session cookie. What changed was not the target. What changed was who was willing to look at the boring part of the code.

What you get out of this: why the most serious bug of the week was ordinary implementation work, and what that means for the scope of your next test.

What happened

CVE-2026-61500 was published on July 13, 2026 and updated on October 1. It affects Rejetto HTTP File Server (HFS) from version 3.0.0 through 3.2.0. The official CVE record states the problem in a sentence that needs no interpretation: HFS derives its session cookie signing key from a non-cryptographic pseudo-random generator and discloses outputs of that same generator to unauthenticated clients during login. With that, a remote attacker can recover the key and forge a valid session cookie. The result is full administrative access and remote code execution, using a feature native to the product itself.

The registered CWE is 338, use of a cryptographically weak PRNG. Severity is critical: CVSS 3.1 of 9.8 and CVSS 4.0 of 9.3, network vector, low complexity, no privileges, no user interaction. The worst case for anything listening on a port.

The finder was researcher Zach Hanley at Horizon3, using Anthropic’s Mythos model inside a vulnerability research pipeline. The company joined Project Glasswing in July 2026 and says it has found many critical vulnerabilities with the model since then. Rejetto’s own fix notice, version 3.2.1, credits Hanley and Anthropic Research by name.

The detail I would save for the conversation with your team is this: Horizon3 itself writes that researchers like them, who avoid pursuing cryptographic flaws, had abandoned that kind of finding before. Not for lack of skill. For lack of the mathematical background to close the impact, and because the time to build the exploit did not pay for itself. A model has no such constraint. It does not get tired on the boring part.

The flaw was not about AI

I will say it again on purpose, because this is where the internet takes it wrong: this is a web application bug. No model, no prompt, no obedient agent. It is a cookie signed with a guessable secret in a Node.js application.

The pattern is everywhere. A team that needs a quick identifier calls Math.random(), drops the result into a token and treats it as a secret. A team that needs randomness for a password or a cookie uses UUID.randomUUID() and considers the matter settled, when that is a different use case. A team that borrowed a session library from a tutorial does not know where the key came from. None of that is innovation. It is Tuesday implementation.

The model’s finding was impressive because it did not stop at the weak PRNG. It connected the weak generator to the code path that leaks outputs from that same generator, worked out whether the leak was enough to recover the key, built the working exploit and ran it. That is not magic. It is sequence. And sequence is exactly what a short human review, done in two hours at the end of the sprint, does not do.

Three months in the wild before anyone used it

The CVE was published on July 13. Real-world exploitation showed up on October 1, according to The Register. Almost a three-month window.

The day after the discovery, the same report says, the CVE was already being exploited. VulnCheck researcher Patrick Garrity said his canaries detected an actor in China targeting genuinely vulnerable hosts in the United States, with overnight activity pointed at servers in the US and Japan, and later hits arriving through a proxy.

That three-month gap is the most useful number in the story, and it is not about AI. It is about inventory. HFS had been in CISA’s known exploited vulnerabilities catalog since 2024 for a template injection flaw that also led to code execution, with known ransomware use. The product was already known. What was missing was somebody looking at the second problem in the same file.

What this changes about your test scope

The pentest I sell is not a zero-day hunt. It is the opposite: picking up the class of bug your team already understands in theory and has never tested in practice, and proving it with a reproducible exploit before somebody else proves it for you.

Three things I would ask for on any Node.js application or anything with a custom session layer.

First, test the unpredictability of session tokens. Cookies and tokens have to be impossible to forge with what an attacker already knows. If your application has a hand-rolled session layer, it is a natural candidate for this class of bug.

Second, administrative functionality is attack surface. HFS has an administrative API that accepts executable code through configuration. Large administrative power that is easy to reach is what turns an authentication problem into an incident.

Third, independent validation with proof. A finding without a running exploit is an opinion. A finding with request, response and observable effect is what closes an audit report.

And honesty cuts the other way too: no pentest finds a leaked cloud key in a repository or a flaw inside an appliance binary. If your risk lives in CI, in a supplier, in patching, that is a different job with a different name. Pentest covers the application you write and run. Widening the scope to everything is the talk of anyone who would rather not say what they do not do.

What to do now

  • If you run Rejetto HFS 3.x, upgrade to 3.2.1 or later. It is the most urgent line in this post.
  • Review where your application draws a secret. If the answer is Math.random(), that is not a secret being drawn.
  • Ask your test supplier for reproducible exploit evidence on authentication findings. Without it, it is a report.
  • Put “what was not tested” into the conversation with your auditor. Declared scope is worth more than presumed coverage.

What stayed with me in this case was not the model. It was that the target was a product already sitting in the CISA catalog, with an earlier flaw of the same class, and nobody had looked. That will keep happening in any codebase until somebody actually tests, with enough time to reach the boring part.

AI is not going to replace whoever decides what to test. It makes whatever you decide late, late enough for the attacker.

Keep reading

Back to blog