Be smart from the start

AI Bug Bounty: Why Human Hackers Still Find What AI Misses

Written by Luca Manara | Aug 25, 2026, 7:06:36 AM

Key takeaways

  • 90% of researchers in the UNGUESS Security community use AI daily, but almost none use it to generate vulnerability hypotheses.
  • AI's two recurring failures: it misses business logic context and it hallucinates vulnerabilities.
  • AI offensive tooling is priced per token, so you pay for process. Bug bounty is priced per validated finding: no finding, no invoice.
  • NIS2, DORA and the Cyber Resilience Act turn human validation and audit trails into a legal requirement, not a preference.

 

Since late 2025 a new class of frontier models, Claude Mythos, GPT Cyber and their peers, has pushed AI vulnerability discovery to the top of every CISO's agenda. The question everyone is asking is the same one: will AI make ethical hackers and bug bounty programs obsolete?

We saw many claims, from what I know still not verified by independent researchers, that these models are able to find large numbers of zero-day vulnerabilities in a seasoned code base. And this is true. And is where AI is very capable: scanning large amounts of code.

By the way, the same models that help you find latent flaws before release will, inevitably, end up in the hands of the people trying to break in.

A whole category of AI-native attacks are coming, leaving CISOs with two practical questions:

How do we actually capture the benefit of all this?

And what’s the most efficient way to do it in money, time, and risk?

We asked researchers from the UNGUESS Security community how they use AI in everyday jobs. This helped us to picture the real situation (at the end, they replicate what a criminal would do).

 

The survey: how ethical hackers actually use AI (and what they don't use it for) 

We asked the community of ethical hackers how they use AI in their everyday job. Not surprisingly, 90% of researchers use AI regularly and intensely. But here is the interesting part: across the eight activity families we mapped, vulnerability hypothesis generation is the one where AI is used least of all. 

Among these families:

Reconnaissance & asset discovery, Understanding unfamiliar code or tech, Vulnerability hypothesis generation, Payload/exploit generation, Writing or refining PoCs, Report writing, Learning/upskilling and Automating repetitive tasks.

They use AI mostly for Understanding unfamiliar code or tech, Writing or refining PoCs and Report writing way less in Reconnaissance & asset discovery and in particular very few use AI for Vulnerability hypothesis generation. Because it is the place where creativity and connecting the dots is important.

Problems from AI usually are on Misses business-logic context and Hallucinated / false vulnerabilities.

Hackers mostly say that AI will raise the bar: easy bugs vanish, only hard bugs pay. They think that AI will reduce earning opportunities for hunters but also that autonomous agents will not compete directly with human hunters.

Counterintuitively they are neutral on this question “Would strong AI-native tooling make you more likely to hunt on a platform?”

AI vulnerability discovery: the problem isn't too few findings, it's too many 

Let’s start with the obvious effect: AI lowers the cost of findings. Frontier models read enormous codebases, help to know how to create a payload, help to chain things together, and increasingly produce working exploits.

As that capability spreads, to your team, to vendors, to attackers, the volume of reported vulnerabilities climbs sharply, even for products that have been pentested for years. The number of CVEs increases everyday more.

 

 That sounds like progress, and partly it is. But most security teams passed the limit of fix everything we find a long time ago. A longer list only helps if each item arrives with reliable context: is it real, is it reachable, does it matter here? Without that, you’ve simply grown the backlog.

Pure-AI offensive tooling struggles on this last mile:

Validation and context. Models are very good at flagging things that resemble vulnerabilities. They are far less dependable at confirming a flaw is genuinely exploitable in your specific environment. Which, today, means a lot of false positives. And false positives quietly burn your most expensive resource: analyst and tech time.

Unpredictable cost. Token-based pricing scales with how much the system runs, not with what it finds. Point a model at a large codebase for testing and the bill grows whether or not anything useful comes back. Traditional pentesting has the same flaw, but on-demand AI makes overruns easier to trigger and harder to cap.

Explainability and control. Most scanning agents still can’t give a clear, auditable account of what they did and didn’t touch, so coverage is hard to guarantee. Worse, some agentic tools take actions on the systems they’re aimed at that you didn’t anticipate, and wouldn’t have authorised.

Sovereignty. Handing a third-party AI deep access to source code and sensitive systems isn’t always legally or politically acceptable; especially for regulated sectors and European organisations operating under stricter data-protection and supply-chain rules.

None of this means the upside is fake. It means the operational stakes are real, and that getting value out of AI offence requires a delivery model that filters signal from noise rather than amplifying both.

 

CISOs don’t want AI. They want outcomes.

Step away from the tooling for a second. When a security leader says I want to use AI for vulnerability discovery, what they almost always mean is: more coverage, faster, at a lower and more predictable cost per real finding, with higher confidence in each one.

That’s an outcome. The frontier model is just one possible means to it – and not necessarily the one you should be buying directly.

The distinction matters because it puts the spotlight on cost structure. Frontier-model APIs and native AI offensive platforms price the way traditional pentesting does: you pay for the process, not the result. You pay for the scan, the tokens, the engagement – including every false positive and every minute of misdirected activity when the configuration is slightly off.

Bug bounty turns that upside down.

 

Bug bounty vs pentest vs AI tooling: pay for results, not activity 

In a bug bounty program you pay when a researcher delivers a validated, exploitable vulnerability – and only then. No finding, no invoice. Cost-per-finding doesn’t swing wildly, because there’s no cost at all when there’s nothing to report.

And you can teach hunters what is important for your context. So that every vulnerability is a real complex attack chain, real exploitable and something to prioritize.

That’s been bug bounty’s differences over traditional pentesting for years. The newly interesting comparison is against AI offensive tooling – and the same economics hold.

 

The researchers already have the AI

Here’s the part the “AI replaces bug bounty” narrative misses entirely: bug-bounty researchers are among the fastest and most ruthless adopters of new technology anywhere in security. Frontier models are already standard equipment in their toolkits: used daily to accelerate reconnaissance, automate the repetitive work, run scans at scale, and probe black-box targets. Each researcher pairs that with their own methodology, intuition, prompting, and frequently their own custom-built tooling.

So the choice was never AI or bug bounty. It’s “yes, and.”

Run a program gives access to independent agentic pentesters, each running a different AI stack against your scope, from a different angle, with different expertise. You get the full reach of frontier models and AI-enabled tools, with two things no raw model gives you:

A human validates every finding. The researcher who submits it stakes their reputation on it, and an expert triage team independently confirms exploitability and assesses real-world risk before anything reaches your team.

You only pay for results. No finding, no fee. Outcomes, not activity.

This is also where the less flattering side of the AI boom gets handled. The same tools that help skilled researchers also flood platforms with low-quality, machine-generated submissions. AI slop that looks plausible and wastes everyone’s time. Volume is up; so is the share of it that’s noise. In that environment, human triage and a vetted researcher community are the thing standing between you and a backlog of confident-sounding garbage.

 

NIS2, DORA and the Cyber Resilience Act: why this matters more in Europe 

For European organisations the calculus is sharper still. NIS2, DORA and the Cyber Resilience Act are converting “good security hygiene” into legal obligation and the CRA in particular requires manufacturers to operate coordinated vulnerability disclosure and handle reports against a clock.

A delivery model that produces validated findings, keeps a clear audit trail, and doesn’t require handing your source code to an opaque foreign AI service it’s increasingly the only model that fits the regulatory and sovereignty constraints European boards are now accountable for.

 

Bug bounty is more relevant in the AI era, not less

Frontier models change the volume and sophistication of what can be found. They don’t change the actual job facing a security team: separating signal from noise, proving exploitability, prioritising the fixes that matter, and not drowning in the process.

A bug bounty program means you pay only for validated vulnerabilities while researchers point their own AI arsenals at your scope, capturing the value of the very latest models and tools without having to run, govern, or pay for any of them yourself.

You wanted the outcomes AI promised. Bug bounty has been delivering them all along. AI just made it better at the job.

 

Do you want to see what a hacker sees? Try our new app, built from the hacker experience: https://security.unguess.io/app

 

FAQ