Lint your LLM eval set. Reliability, items scored against a wrong answer, near-duplicates, and how many test cases you can drop without changing the ranking. Zero dependencies.

ai benchmark benchmarking cli data-quality eval evals evaluation item-analysis llm llm-evaluation llmops machine-learning psychometrics python reliability statistics test-quality
1 Open Issue Need Help Last updated: Aug 11, 2026

Open Issues Need Help

View All on GitHub
help wanted question

Lint your LLM eval set. Reliability, items scored against a wrong answer, near-duplicates, and how many test cases you can drop without changing the ranking. Zero dependencies.

Python
#ai#benchmark#benchmarking#cli#data-quality#eval#evals#evaluation#item-analysis#llm#llm-evaluation#llmops#machine-learning#psychometrics#python#reliability#statistics#test-quality