Practical evaluation techniques, benchmarks, and examples for production AI agents: RAG, tool use, planning, memory, safety, reliability, and cost.

6 stars 7 forks 6 watchers Jupyter Notebook MIT License
agent-evaluation agentic-ai ai-agents benchmark evaluation genai langgraph llm llm-evaluation openai rag tool-calling
8 Open Issues Need Help Last updated: Aug 13, 2026

Open Issues Need Help

View All on GitHub
enhancement help wanted evaluation

Practical evaluation techniques, benchmarks, and examples for production AI agents: RAG, tool use, planning, memory, safety, reliability, and cost.

Jupyter Notebook
#agent-evaluation#agentic-ai#ai-agents#benchmark#evaluation#genai#langgraph#llm#llm-evaluation#openai#rag#tool-calling
enhancement good first issue evaluation

Practical evaluation techniques, benchmarks, and examples for production AI agents: RAG, tool use, planning, memory, safety, reliability, and cost.

Jupyter Notebook
#agent-evaluation#agentic-ai#ai-agents#benchmark#evaluation#genai#langgraph#llm#llm-evaluation#openai#rag#tool-calling
Latency Evaluation about 3 hours ago
enhancement good first issue evaluation

Practical evaluation techniques, benchmarks, and examples for production AI agents: RAG, tool use, planning, memory, safety, reliability, and cost.

Jupyter Notebook
#agent-evaluation#agentic-ai#ai-agents#benchmark#evaluation#genai#langgraph#llm#llm-evaluation#openai#rag#tool-calling
Cost Evaluation about 3 hours ago
enhancement good first issue evaluation

Practical evaluation techniques, benchmarks, and examples for production AI agents: RAG, tool use, planning, memory, safety, reliability, and cost.

Jupyter Notebook
#agent-evaluation#agentic-ai#ai-agents#benchmark#evaluation#genai#langgraph#llm#llm-evaluation#openai#rag#tool-calling
enhancement help wanted benchmark evaluation

Practical evaluation techniques, benchmarks, and examples for production AI agents: RAG, tool use, planning, memory, safety, reliability, and cost.

Jupyter Notebook
#agent-evaluation#agentic-ai#ai-agents#benchmark#evaluation#genai#langgraph#llm#llm-evaluation#openai#rag#tool-calling
enhancement help wanted benchmark evaluation

Practical evaluation techniques, benchmarks, and examples for production AI agents: RAG, tool use, planning, memory, safety, reliability, and cost.

Jupyter Notebook
#agent-evaluation#agentic-ai#ai-agents#benchmark#evaluation#genai#langgraph#llm#llm-evaluation#openai#rag#tool-calling
enhancement help wanted evaluation

Practical evaluation techniques, benchmarks, and examples for production AI agents: RAG, tool use, planning, memory, safety, reliability, and cost.

Jupyter Notebook
#agent-evaluation#agentic-ai#ai-agents#benchmark#evaluation#genai#langgraph#llm#llm-evaluation#openai#rag#tool-calling
Tool Calling Benchmark about 3 hours ago
enhancement help wanted benchmark evaluation

Practical evaluation techniques, benchmarks, and examples for production AI agents: RAG, tool use, planning, memory, safety, reliability, and cost.

Jupyter Notebook
#agent-evaluation#agentic-ai#ai-agents#benchmark#evaluation#genai#langgraph#llm#llm-evaluation#openai#rag#tool-calling