Why AI Vulnerability Discovery is Overwhelming DevSecOps Teams
- When Anthropic introduced Project Glasswing in April, the tech industry braced for impact.
- Garrity and Vulncheck tracked vulnerabilities discovered with AI against data from the Berkeley Vulnerability Research Initiative and Vulncheck's known-exploited vulnerability database.
- Anthropic reported 26,153 total findings through the initiative, with 2,736 or 10.5% reaching the ledger.
Anthropic’s Project Glasswing Initiative Faces Human Triage Bottlenecks in AI-Assisted Vulnerability Discovery
When Anthropic introduced Project Glasswing in April, the tech industry braced for impact. Anthropic updated its Vulnerability Disclosure ledger for the first time in August. That’s roughly 1%. “You discover a vulnerability — that doesn’t mean it’s necessarily useful for an attacker. It doesn’t mean it’s necessarily going to be exploitable with certain conditions,” Garrity said. It isn’t. “No one is talking about… the risks [that arise] when you deploy these technologies,” Garrity said. “The bigger shift I’m seeing is actually the attack surface expanding substantially. AI products themselves are being targeted and hit very fast.” Glasswing results: humans ‘rate-limiting’? “The Glasswing approach came out and was like, everything is critical and high. What we’re finding is the maintainers are scoring things not all critical and high,” said Garrity. “This wasn’t their subject domain.”
Vulnerability Volume Versus Actual Exploitation
Garrity and Vulncheck tracked vulnerabilities discovered with AI against data from the Berkeley Vulnerability Research Initiative and Vulncheck’s known-exploited vulnerability database. Out of 1,061 total vulnerabilities attributed to AI-assisted discovery, researchers found that only 14 had been listed as exploited in the wild as of June 2026, representing roughly 1%.
You discover a vulnerability — that doesn’t mean it’s necessarily useful for an attacker. It doesn’t mean it’s necessarily going to be exploitable with certain conditions,
Garrity said.
Anthropic reported 26,153 total findings through the initiative, with 2,736 or 10.5% reaching the ledger. Software maintainers received 2,096 reports, representing 8% of the total findings, while 245 findings were withdrawn and 202, or 0.8%, were marked fixed.

Human Triage Bottlenecks and Severity Discrepancies
Software security experts point to human triage as the primary rate-limiting step where engineering teams must determine what findings matter, evaluate severity, and decide whether and how to apply fixes. Anthropic’s ledger explicitly described independent human triage as a rate-limiting step.
There was so much agita around [the] ‘Vulnpocalypse’, around everybody being able to find new zero-day vulnerabilities and exploit them, the impact that was going to have,
said Neil Carpenter, an independent security evangelist. And then the secondary piece of, ‘How do we fix so many vulnerabilities?’ It becomes a burden on defenders, both developers and blue teams downstream.
Analysis of the ledger also revealed a significant gap between how the AI model rated vulnerability severity compared to human maintainers. Claude assessed 91.5% of its findings as high or critical severity, whereas project maintainers classified only 51.3% of those same findings in the same upper tiers.
In May, Daniel Stenberg, the founder and lead developer of the Linux curl command-line utility project, detailed how Mythos identified five confirmed vulnerabilities in the software. Stenberg noted that his team reduced that count to a single low-severity CVE scheduled for release in version 8.21.0, dismissing the other four findings as three documentation-related false positives and one ordinary bug.
The Glasswing approach came out and was like, everything is critical and high. What we’re finding is the maintainers are scoring things not all critical and high,
Garrity noted, adding that the AI model development team may have lacked specific subject domain expertise to guide Mythos accurately.
Engineering Workflows and New Attack Surfaces
While organizations can deploy automation to parse accumulated vulnerability data, security architects emphasize that artificial intelligence does not eliminate the need for human oversight. Michele Chubirka, a senior principal security architect who participated in Project Glasswing at Red Hat, helped build a custom workflow engine combining harnesses and deterministic security tools to scan OpenShift product repositories.
People think that AI is magic, that a frontier model is going to magically do this stuff for you. It isn’t,
Chubirka said, speaking in a personal capacity.
Deploying AI tools directly into production workflows expands the overall security attack surface. Garrity observed that AI products themselves are actively targeted and exploited at a rapid pace, introducing entirely new software bugs and operational risks that security teams must manage.
