Open Issues Need Help
View All on GitHub Default the judge to temperature 0 (and warn when it isn't) about 2 hours ago
enhancement good first issue
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
Ruby
#anthropic#evaluation-framework#evaluation-metrics#llm#llm-as-judge#llm-eval#llm-evaluation#llm-evaluation-framework#llm-evaluation-metrics#llmops#mcp#ollama#openai#prompt-engineering#prompt-testing#rails#rails-engine#ruby#ruby-on-rails