Agents claim work is done that isn't. This makes them prove it: evidence-gated turns, self-testing gates, live-state checks, and a ledger that catches "all done" lies. One-command Claude Code integration.

agents ai-agents claude-code evals guardrails llm python quality-assurance reliability testing
2 Open Issues Need Help Last updated: Aug 7, 2026

Open Issues Need Help

View All on GitHub
help wanted good first issue

Agents claim work is done that isn't. This makes them prove it: evidence-gated turns, self-testing gates, live-state checks, and a ledger that catches "all done" lies. One-command Claude Code integration.

Python
#agents#ai-agents#claude-code#evals#guardrails#llm#python#quality-assurance#reliability#testing
help wanted good first issue

Agents claim work is done that isn't. This makes them prove it: evidence-gated turns, self-testing gates, live-state checks, and a ledger that catches "all done" lies. One-command Claude Code integration.

Python
#agents#ai-agents#claude-code#evals#guardrails#llm#python#quality-assurance#reliability#testing