LLM security analyzer rediscovers 68% of real-world CVEs without frontier models
Researchers have built HoF-Bench, a rigorous evaluation framework grounded in 95 real CVEs discovered by AISLE, an LLM-based security analyzer that has already identified over 280 vulnerabilities in production systems like OpenSSL and curl. The benchmark enforces strict validation: detectors must pinpoint the exact code path, root cause, and attack surface without access to CVE metadata or frontier models. A minimal LLM analyzer recovered 68 percent of vulnerabilities under this protocol, establishing a reproducible standard for measuring AI-driven security research and challenging claims that only large frontier models can perform meaningful vulnerability discovery.
Modelwire context
Skeptical readThe paper establishes a reproducible evaluation protocol, but the actual finding is more modest than the framing suggests: a minimal LLM recovered 68 percent of vulnerabilities, not all of them. The implicit claim that frontier models are unnecessary for vulnerability discovery conflicts with the fact that AISLE (which discovered the 95 CVEs in the benchmark) was itself a scaled system.
This connects to the broader pattern from 'Lottery Tickets Are Not Deployment Tickets' (late July), which showed that lab-validated optimizations often fail in production. Here, the benchmark proves that smaller models can match frontier performance on a curated task, but doesn't address whether they generalize to undiscovered vulnerabilities or whether the 32 percent gap matters for real security workflows. The rigor is genuine, but the practical implication remains unclear.
If AISLE or a similar minimal-model system discovers novel CVEs in the next six months that pass the HoF-Bench validation protocol without human annotation, that confirms the benchmark's predictive value. If instead the 68 percent recovery rate holds only for the 95 known CVEs and new discoveries require larger models, the claim collapses.
Coverage we drew on
- Lottery Tickets Are Not Deployment Tickets · arXiv cs.LG
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsAISLE · HoF-Bench · OpenSSL · curl · GnuTLS
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.LG originally reported this story as “HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.