Attribute eval-score movement to the system or the LLM judge using a frozen human anchor set and kappa.

ai-engineering ci drift-detection evals llm-as-judge llm-evaluation python reliability
1 Open Issue Need Help Last updated: Aug 5, 2026

Open Issues Need Help

View All on GitHub

Attribute eval-score movement to the system or the LLM judge using a frozen human anchor set and kappa.

Python
#ai-engineering#ci#drift-detection#evals#llm-as-judge#llm-evaluation#python#reliability