A sealed capability-ladder benchmark for LLM-driven vulnerability reproduction: 68 real zero-day bugs across 40 open-source projects (C/C++/Java), graded by a deterministic remote oracle — no answer key ships.

ai-agents benchmark fuzzing llm security vulnerability-detection
2 Open Issues Need Help Last updated: Jul 30, 2026

Open Issues Need Help

View All on GitHub

A sealed capability-ladder benchmark for LLM-driven vulnerability reproduction: 68 real zero-day bugs across 40 open-source projects (C/C++/Java), graded by a deterministic remote oracle — no answer key ships.

Python
#ai-agents#benchmark#fuzzing#llm#security#vulnerability-detection

A sealed capability-ladder benchmark for LLM-driven vulnerability reproduction: 68 real zero-day bugs across 40 open-source projects (C/C++/Java), graded by a deterministic remote oracle — no answer key ships.

Python
#ai-agents#benchmark#fuzzing#llm#security#vulnerability-detection