LLM-as-judge evaluation harness where the judge itself must pass a Cohen's-kappa calibration gate

ai-governance anthropic cohens-kappa evals inter-rater-reliability llm-as-a-judge llm-evaluation python
3 Open Issues Need Help Last updated: Aug 28, 2026

Open Issues Need Help

View All on GitHub

LLM-as-judge evaluation harness where the judge itself must pass a Cohen's-kappa calibration gate

Python
#ai-governance#anthropic#cohens-kappa#evals#inter-rater-reliability#llm-as-a-judge#llm-evaluation#python
help wanted good first issue

LLM-as-judge evaluation harness where the judge itself must pass a Cohen's-kappa calibration gate

Python
#ai-governance#anthropic#cohens-kappa#evals#inter-rater-reliability#llm-as-a-judge#llm-evaluation#python

LLM-as-judge evaluation harness where the judge itself must pass a Cohen's-kappa calibration gate

Python
#ai-governance#anthropic#cohens-kappa#evals#inter-rater-reliability#llm-as-a-judge#llm-evaluation#python