Mythos Vulnerability Firehose Hits a Human Bottleneck
A new analysis of public data from Anthropic's Project Glasswing has highlighted a significant gap between the number of vulnerability findings generated by its Claude frontier model and those that ultimately prove to be real, serious, and worth fixing.
A new analysis of public data from Anthropic's Project Glasswing has highlighted a significant gap between the number of vulnerability findings generated by its Claude frontier model and those that ultimately prove to be real, serious, and worth fixing. The distinction matters because it suggests that the bottleneck in vulnerability research may increasingly lie in validating new flaws and coordinating their remediation rather than in discovering them. Patrick Garrity, a security researcher at VulnCheck, recently analyzed Anthropic’s Vulnerability Disclosure Ledger , which is a public record tracking Project Glasswing-related findings as they move through the vulnerability disclosure and remediation process. The analysis showed that Anthropic's Claude Mythos generated a total of 26,153 vulnerability findings across numerous software projects since Project Glasswing's launch in April 2026.
But it "is much different than the narrative frontier model providers have positioned," which has largely focused on AI's ability to dramatically accelerate vulnerability discovery. Garrity's analysis also raised questions about Anthropic's claims regarding the accuracy of Mythos' vulnerability findings and the model's ability to assess their severity. Related: AI's Vulnerability Surge May Be More Manageable Than First Feared "My gut tells me the team didn't prompt Claude with detailed instructions on how to determine severity, or if they did, it wasn't well thought out, resulting in higher severity determinations," Garrity says. Related: Attackers Pounce on Critical Artifactory Bug Following Disclosure The disparity highlights a potential shift in the economics of vulnerability research , Williams says.
But it "is much different than the narrative frontier model providers have positioned," which has largely focused on AI's ability to dramatically accelerate vulnerability discovery. Garrity's analysis also raised questions about Anthropic's claims regarding the accuracy of Mythos' vulnerability findings and the model's ability to assess their severity. Related: AI's Vulnerability Surge May Be More Manageable Than First Feared "My gut tells me the team didn't prompt Claude with detailed instructions on how to determine severity, or if they did, it wasn't well thought out, resulting in higher severity determinations," Garrity says. Related: Attackers Pounce on Critical Artifactory Bug Following Disclosure The disparity highlights a potential shift in the economics of vulnerability research , Williams says.
What changed
He noted that the 202 findings marked as fixed in the vulnerability ledger are notably fewer than the 245 vulnerabilities marked as withdrawn.
"Scanning a 2-million-line codebase cost around $315 in API charges and triaging the results cost around $128,000." VulnCheck's analysis comes as Project Glasswing and similar efforts to apply AI to vulnerability discovery across the industry have begun producing a massive volume of newly identified flaws.
Who is affected
"I'd like to see the CVSS metrics used to generate severity and the CWEs used to help better understand the actual weaknesses, both of which are industry standards expected when disclosing vulnerabilities." The questions around AI's ability to accurately assess vulnerability findings are not unique to Glasswing.
Microsoft's record-setting September Patch Tuesday this week, which addressed 974 vulnerabilities, offers an example of the scale of the challenge that organizations face as AI accelerates vulnerability discovery.
The technical picture
But it "is much different than the narrative frontier model providers have positioned," which has largely focused on AI's ability to dramatically accelerate vulnerability discovery.
Garrity's analysis also raised questions about Anthropic's claims regarding the accuracy of Mythos' vulnerability findings and the model's ability to assess their severity.
Related: AI's Vulnerability Surge May Be More Manageable Than First Feared "My gut tells me the team didn't prompt Claude with detailed instructions on how to determine severity, or if they did, it wasn't well thought out, resulting in higher severity determinations," Garrity says.
Related: Attackers Pounce on Critical Artifactory Bug Following Disclosure The disparity highlights a potential shift in the economics of vulnerability research , Williams says.
The trend is creating new challenges for security and vulnerability remediation teams that have to validate, prioritize, and patch those flaws.
What to watch next
Watch for revised fixed-version guidance and confirmation that mitigations are holding in affected environments.
What remains unknown
The available reporting does not establish whether the issue is being actively exploited in the wild.
The available reporting does not establish who is behind the activity, if an attacker is involved.