Open Issues Need Help
View All on GitHub vLLM collector: KV cache, queue depth, and token throughput about 2 hours ago
help wanted collector
Your GPU is not 90% busy. truthscale measures what each node can actually deliver, reports the gap, and makes that the signal you scale on.
Go
#autoscaling#cloud-native#dcgm#finops#go#gpu#gpu-monitoring#gpu-utilization#inference-serving#kubernetes#kv-cache#llm-inference#mlops#nvidia#observability#platform-engineering#prometheus#rke2#sre#vllm
DCGM collector: read the real fields, and report what the driver will not give you about 2 hours ago
help wanted collector needs-gpu
Your GPU is not 90% busy. truthscale measures what each node can actually deliver, reports the gap, and makes that the signal you scale on.
Go
#autoscaling#cloud-native#dcgm#finops#go#gpu#gpu-monitoring#gpu-utilization#inference-serving#kubernetes#kv-cache#llm-inference#mlops#nvidia#observability#platform-engineering#prometheus#rke2#sre#vllm