Open Issues Need Help
View All on GitHub qwen3: PerToken decode-graph memory grows linearly with batch bucket and is not budgeted about 3 hours ago
help wanted qwen3 hw:1-gpu
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Rust
#cuda#cuda-kernels#deepseek#gpu#inference#inference-engine#kimi#kimi-k2#kv-cache#llm#llm-inference#llm-serving#model-serving#moe#openai-api#paged-attention#qwen#qwen3#rust#vllm