Open Issues Need Help
View All on GitHub Create a task/experiment init command about 13 hours ago
good first issue
A framework for evaluating AI coding agents and their skills with sandboxing, reproducibility, and data-driven analysis.
Python
#agent-evaluation#agentic-ai#ai-agents#benchmark#benchmarking#claude#claude-code#claude-code-skill#claude-skills#cli#codex#coding-agents#coding-agents-evaluation#evals#evaluation-framework#llm-evaluation#skills#skills-evaluation