Open Issues Need Help
View All on GitHub #2 priority task: Improve the scoring system, introduce the difficulty coefficient for each challenge about 1 hour ago
help wanted
A sealed capability-ladder benchmark for LLM-driven vulnerability reproduction: 68 real zero-day bugs across 40 open-source projects (C/C++/Java), graded by a deterministic remote oracle — no answer key ships.
Python
#ai-agents#benchmark#fuzzing#llm#security#vulnerability-detection
help wanted
A sealed capability-ladder benchmark for LLM-driven vulnerability reproduction: 68 real zero-day bugs across 40 open-source projects (C/C++/Java), graded by a deterministic remote oracle — no answer key ships.
Python
#ai-agents#benchmark#fuzzing#llm#security#vulnerability-detection