Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.

anthropic evaluation-framework evaluation-metrics llm llm-as-judge llm-eval llm-evaluation llm-evaluation-framework llm-evaluation-metrics llmops mcp ollama openai prompt-engineering prompt-testing rails rails-engine ruby ruby-on-rails
1 Open Issue Need Help Last updated: Jul 28, 2026

Open Issues Need Help

View All on GitHub
enhancement good first issue

Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.

Ruby
#anthropic#evaluation-framework#evaluation-metrics#llm#llm-as-judge#llm-eval#llm-evaluation#llm-evaluation-framework#llm-evaluation-metrics#llmops#mcp#ollama#openai#prompt-engineering#prompt-testing#rails#rails-engine#ruby#ruby-on-rails