Framework-native, reproducible prompt-injection benchmark, CI gate, ImpactTwin procurement tests, and attested ProofRun evidence for AI agents.

agent-benchmark agentdojo agentic-ai ai-agents ai-safety autogen crewai dspy github-actions langchain leaderboard llm-evaluation llm-security openai-agents procurement prompt-injection pydantic-ai red-teaming reproducible-research security-benchmark
1 Open Issue Need Help Last updated: Aug 13, 2026

Open Issues Need Help

View All on GitHub
enhancement help wanted

Framework-native, reproducible prompt-injection benchmark, CI gate, ImpactTwin procurement tests, and attested ProofRun evidence for AI agents.

Python
#agent-benchmark#agentdojo#agentic-ai#ai-agents#ai-safety#autogen#crewai#dspy#github-actions#langchain#leaderboard#llm-evaluation#llm-security#openai-agents#procurement#prompt-injection#pydantic-ai#red-teaming#reproducible-research#security-benchmark