Open Issues Need Help
View All on GitHub good first issue
Attribute eval-score movement to the system or the LLM judge using a frozen human anchor set and kappa.
Python
#ai-engineering#ci#drift-detection#evals#llm-as-judge#llm-evaluation#python#reliability