Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

1 stars 0 forks 1 watchers Python Apache License 2.0
agent-evaluation agent-skills agentic-ai ai-agents benchmarking codex coding-agents evals llm-agents llm-evaluation reproducibility software-engineering
7 Open Issues Need Help Last updated: Aug 14, 2026

Open Issues Need Help

View All on GitHub
help wanted research needs-design difficulty:flagship security-design

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python
#agent-evaluation#agent-skills#agentic-ai#ai-agents#benchmarking#codex#coding-agents#evals#llm-agents#llm-evaluation#reproducibility#software-engineering
help wanted research replication needs-design difficulty:flagship

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python
#agent-evaluation#agent-skills#agentic-ai#ai-agents#benchmarking#codex#coding-agents#evals#llm-agents#llm-evaluation#reproducibility#software-engineering
enhancement help wanted community needs-design difficulty:advanced

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python
#agent-evaluation#agent-skills#agentic-ai#ai-agents#benchmarking#codex#coding-agents#evals#llm-agents#llm-evaluation#reproducibility#software-engineering
help wanted runner-adapter needs-design difficulty:advanced

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python
#agent-evaluation#agent-skills#agentic-ai#ai-agents#benchmarking#codex#coding-agents#evals#llm-agents#llm-evaluation#reproducibility#software-engineering
help wanted python needs-design difficulty:advanced

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python
#agent-evaluation#agent-skills#agentic-ai#ai-agents#benchmarking#codex#coding-agents#evals#llm-agents#llm-evaluation#reproducibility#software-engineering
good first issue help wanted task-pack

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python
#agent-evaluation#agent-skills#agentic-ai#ai-agents#benchmarking#codex#coding-agents#evals#llm-agents#llm-evaluation#reproducibility#software-engineering
documentation good first issue community

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python
#agent-evaluation#agent-skills#agentic-ai#ai-agents#benchmarking#codex#coding-agents#evals#llm-agents#llm-evaluation#reproducibility#software-engineering