Your agent said Done. nuhuh runs the experiment with fresh tests, real exit codes and actual files, then hands back a receipt it wrote, not one your agent dictated.

1 stars 0 forks 1 watchers TypeScript MIT License
ai-agents benchmark claude-code codex stop-hook testing verification
3 Open Issues Need Help Last updated: Aug 11, 2026

Open Issues Need Help

View All on GitHub

Your agent said Done. nuhuh runs the experiment with fresh tests, real exit codes and actual files, then hands back a receipt it wrote, not one your agent dictated.

TypeScript
#ai-agents#benchmark#claude-code#codex#stop-hook#testing#verification

Your agent said Done. nuhuh runs the experiment with fresh tests, real exit codes and actual files, then hands back a receipt it wrote, not one your agent dictated.

TypeScript
#ai-agents#benchmark#claude-code#codex#stop-hook#testing#verification
good first issue

Your agent said Done. nuhuh runs the experiment with fresh tests, real exit codes and actual files, then hands back a receipt it wrote, not one your agent dictated.

TypeScript
#ai-agents#benchmark#claude-code#codex#stop-hook#testing#verification