Open Issues Need Help
View All on GitHub good first issue
Diagnose whether a low-kappa LLM judge panel fails from item ambiguity or rubric underspecification.
Python
#ai-evaluation#bias-detection#cohen-kappa#evaluation-metrics#fleiss-kappa#inter-rater-reliability#llm-as-a-judge#llm-evaluation#python#reliability