Open Issues Need Help
View All on GitHubBenchmark how well any model performs real-world, long-horizon coding tasks across coding agents (harnesses) -- Claude Code and pi. Run models three ways -- Anthropic on Amazon Bedrock, open-weight on Bedrock via a LiteLLM proxy, or self-hosted on EC2 with vLLM -- score them with an LLM judge, and plot the cost/quality Pareto frontier.
Benchmark how well any model performs real-world, long-horizon coding tasks across coding agents (harnesses) -- Claude Code and pi. Run models three ways -- Anthropic on Amazon Bedrock, open-weight on Bedrock via a LiteLLM proxy, or self-hosted on EC2 with vLLM -- score them with an LLM judge, and plot the cost/quality Pareto frontier.