Reproducible ML inference benchmarking across heterogeneous hardware: one YAML config drives llama.cpp and ONNX Runtime sweeps on laptops, Raspberry Pi (SSH) and GPUs, with TTFT, throughput, per-rail power, thermal and utilization telemetry in self-describing records, plus publication-quality tables and figures.

benchmarking edge-ai energy-measurement llama-cpp llm-inference onnxruntime raspberry-pi reproducible-research roofline-model ttft
3 Open Issues Need Help Last updated: Aug 24, 2026

Open Issues Need Help

View All on GitHub

Reproducible ML inference benchmarking across heterogeneous hardware: one YAML config drives llama.cpp and ONNX Runtime sweeps on laptops, Raspberry Pi (SSH) and GPUs, with TTFT, throughput, per-rail power, thermal and utilization telemetry in self-describing records, plus publication-quality tables and figures.

Python
#benchmarking#edge-ai#energy-measurement#llama-cpp#llm-inference#onnxruntime#raspberry-pi#reproducible-research#roofline-model#ttft
enhancement help wanted

Reproducible ML inference benchmarking across heterogeneous hardware: one YAML config drives llama.cpp and ONNX Runtime sweeps on laptops, Raspberry Pi (SSH) and GPUs, with TTFT, throughput, per-rail power, thermal and utilization telemetry in self-describing records, plus publication-quality tables and figures.

Python
#benchmarking#edge-ai#energy-measurement#llama-cpp#llm-inference#onnxruntime#raspberry-pi#reproducible-research#roofline-model#ttft
help wanted good first issue

Reproducible ML inference benchmarking across heterogeneous hardware: one YAML config drives llama.cpp and ONNX Runtime sweeps on laptops, Raspberry Pi (SSH) and GPUs, with TTFT, throughput, per-rail power, thermal and utilization telemetry in self-describing records, plus publication-quality tables and figures.

Python
#benchmarking#edge-ai#energy-measurement#llama-cpp#llm-inference#onnxruntime#raspberry-pi#reproducible-research#roofline-model#ttft