Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

1 stars 0 forks 1 watchers Python Apache License 2.0
agent-benchmark agent-evaluation ai-agents ai-safety aiops benchmark evaluation-framework high-performance-computing hpc langfuse llm llm-agents llm-benchmark llm-evaluation mcp python rbac reproducibility slurm tool-use
19 Open Issues Need Help Last updated: Aug 8, 2026

Open Issues Need Help

View All on GitHub
help wanted area: infra effort: medium type: tech-debt

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
help wanted good first issue area: infra effort: small type: tech-debt

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
documentation help wanted

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted area: scorers effort: large

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
help wanted area: reporting effort: large type: epic

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted area: a2a effort: large

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted area: mcp effort: medium

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted area: api effort: large

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted area: api effort: medium

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted area: api effort: medium

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted area: api effort: large

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
documentation help wanted area: docs effort: small

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
bug help wanted area: reporting effort: medium

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
help wanted area: infra effort: large type: tech-debt

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted good first issue area: cli effort: medium

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
documentation help wanted good first issue area: docs effort: medium

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted good first issue area: cli effort: small

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted good first issue area: cli effort: medium

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use
enhancement help wanted good first issue area: cli effort: small

Role-aware, permission-enforced benchmark for evaluating AI agents that operate HPC systems — SLURM, telemetry, RBAC, docs, facility. Trace-based, reproducible, tool-using. 80 tasks × 26 environments × 12 scorers.

Python
#agent-benchmark#agent-evaluation#ai-agents#ai-safety#aiops#benchmark#evaluation-framework#high-performance-computing#hpc#langfuse#llm#llm-agents#llm-benchmark#llm-evaluation#mcp#python#rbac#reproducibility#slurm#tool-use