Open Issues Need Help
View All on GitHub help wanted good first issue
LLM-as-judge evaluation harness where the judge itself must pass a Cohen's-kappa calibration gate
Python
#ai-governance#anthropic#cohens-kappa#evals#inter-rater-reliability#llm-as-a-judge#llm-evaluation#python
Add 10 harder boundary cases to the golden set about 3 hours ago
help wanted good first issue
LLM-as-judge evaluation harness where the judge itself must pass a Cohen's-kappa calibration gate
Python
#ai-governance#anthropic#cohens-kappa#evals#inter-rater-reliability#llm-as-a-judge#llm-evaluation#python
Rewrite Completeness (C) rubric anchors — current weighted kappa is 0.20 about 3 hours ago
help wanted good first issue
LLM-as-judge evaluation harness where the judge itself must pass a Cohen's-kappa calibration gate
Python
#ai-governance#anthropic#cohens-kappa#evals#inter-rater-reliability#llm-as-a-judge#llm-evaluation#python