Agents claim work is done that isn't. This makes them prove it: evidence-gated turns, self-testing gates, live-state checks, and a ledger that catches "all done" lies. One-command Claude Code integration.

agents ai-agents claude-code evals guardrails llm python quality-assurance reliability testing
1 Open Issue Need Help Last updated: Aug 8, 2026

Open Issues Need Help

View All on GitHub

Agents claim work is done that isn't. This makes them prove it: evidence-gated turns, self-testing gates, live-state checks, and a ledger that catches "all done" lies. One-command Claude Code integration.

Python
#agents#ai-agents#claude-code#evals#guardrails#llm#python#quality-assurance#reliability#testing