Open Issues Need Help
View All on GitHub Improve prompt-processing (prefill) throughput about 3 hours ago
enhancement help wanted
Frontier-class LLM inference on a laptop CPU — gpt-oss:20b at ~110 tok/s on Apple M4 Max, 7.5x llama.cpp, no GPU. From-scratch NEON/SME kernels in Rust.
Rust
#apple-silicon#cpu#gpt-oss#inference#llm#mixture-of-experts#neon#quantization#rust#simd
enhancement help wanted
Frontier-class LLM inference on a laptop CPU — gpt-oss:20b at ~110 tok/s on Apple M4 Max, 7.5x llama.cpp, no GPU. From-scratch NEON/SME kernels in Rust.
Rust
#apple-silicon#cpu#gpt-oss#inference#llm#mixture-of-experts#neon#quantization#rust#simd
Port the decode kernels to x86 (AVX-512 / VNNI) about 3 hours ago
enhancement help wanted
Frontier-class LLM inference on a laptop CPU — gpt-oss:20b at ~110 tok/s on Apple M4 Max, 7.5x llama.cpp, no GPU. From-scratch NEON/SME kernels in Rust.
Rust
#apple-silicon#cpu#gpt-oss#inference#llm#mixture-of-experts#neon#quantization#rust#simd