Benchmark how well any model performs real-world, long-horizon coding tasks across coding agents (harnesses) -- Claude Code and pi. Run models three ways -- Anthropic on Amazon Bedrock, open-weight on Bedrock via a LiteLLM proxy, or self-hosted on EC2 with vLLM -- score them with an LLM judge, and plot the cost/quality Pareto frontier.

3 stars 1 forks 3 watchers Python MIT No Attribution
agentic-ai amazon-bedrock benchmark claude-code coding-agent cost-optimization harness llm llm-evaluation open-weight-models pareto-frontier pi self-hosted-llm vllm
2 Open Issues Need Help Last updated: Aug 14, 2026

Open Issues Need Help

View All on GitHub

Benchmark how well any model performs real-world, long-horizon coding tasks across coding agents (harnesses) -- Claude Code and pi. Run models three ways -- Anthropic on Amazon Bedrock, open-weight on Bedrock via a LiteLLM proxy, or self-hosted on EC2 with vLLM -- score them with an LLM judge, and plot the cost/quality Pareto frontier.

Python
#agentic-ai#amazon-bedrock#benchmark#claude-code#coding-agent#cost-optimization#harness#llm#llm-evaluation#open-weight-models#pareto-frontier#pi#self-hosted-llm#vllm

Benchmark how well any model performs real-world, long-horizon coding tasks across coding agents (harnesses) -- Claude Code and pi. Run models three ways -- Anthropic on Amazon Bedrock, open-weight on Bedrock via a LiteLLM proxy, or self-hosted on EC2 with vLLM -- score them with an LLM judge, and plot the cost/quality Pareto frontier.

Python
#agentic-ai#amazon-bedrock#benchmark#claude-code#coding-agent#cost-optimization#harness#llm#llm-evaluation#open-weight-models#pareto-frontier#pi#self-hosted-llm#vllm