Open Issues Need Help
View All on GitHub Feedback wanted: real repeat-run exports about 2 hours ago
help wanted question
Lint your LLM eval set. Reliability, items scored against a wrong answer, near-duplicates, and how many test cases you can drop without changing the ranking. Zero dependencies.
Python
#ai#benchmark#benchmarking#cli#data-quality#eval#evals#evaluation#item-analysis#llm#llm-evaluation#llmops#machine-learning#psychometrics#python#reliability#statistics#test-quality