Open Issues Need Help
View All on GitHub security: prototype containerized evaluator isolation about 3 hours ago
help wanted security
TypeScript service and reproducible RL-style environment for evaluating coding-agent reliability, refactoring, and performance work.
Python
#ai-evaluation#benchmarking#coding-agents#nodejs#python#reinforcement-learning#reliability#software-engineering#testing#typescript
docs: add an end-to-end task-author walkthrough about 3 hours ago
documentation good first issue
TypeScript service and reproducible RL-style environment for evaluating coding-agent reliability, refactoring, and performance work.
Python
#ai-evaluation#benchmarking#coding-agents#nodejs#python#reinforcement-learning#reliability#software-engineering#testing#typescript
test: expand async idempotency and scheduled-delivery edge cases about 3 hours ago
help wanted testing
TypeScript service and reproducible RL-style environment for evaluating coding-agent reliability, refactoring, and performance work.
Python
#ai-evaluation#benchmarking#coding-agents#nodejs#python#reinforcement-learning#reliability#software-engineering#testing#typescript
docs: add Windows and WSL verification guide about 3 hours ago
documentation good first issue
TypeScript service and reproducible RL-style environment for evaluating coding-agent reliability, refactoring, and performance work.
Python
#ai-evaluation#benchmarking#coding-agents#nodejs#python#reinforcement-learning#reliability#software-engineering#testing#typescript