arizqi

arizqi/cpubrrr

Frontier-class LLM inference on a laptop CPU — gpt-oss:20b at ~110 tok/s on Apple M4 Max, 7.5x llama.cpp, no GPU. From-scratch NEON/SME kernels in Rust.

3 good first / help-wanted issues · Rust · last activity Jul 23, 2026

12 stars 2 forks 12 watchers Rust Apache License 2.0
apple-silicon cpu gpt-oss inference llm mixture-of-experts neon quantization rust simd
3 Open Issues Need Help Last updated: Jul 23, 2026

Open Issues Need Help

View All on GitHub
enhancement help wanted good first issue
arizqi/cpubrrr
12

Frontier-class LLM inference on a laptop CPU — gpt-oss:20b at ~110 tok/s on Apple M4 Max, 7.5x llama.cpp, no GPU. From-scratch NEON/SME kernels in Rust.

Rust
#apple-silicon#cpu#gpt-oss#inference#llm#mixture-of-experts#neon#quantization#rust#simd
enhancement help wanted
arizqi/cpubrrr
12

Frontier-class LLM inference on a laptop CPU — gpt-oss:20b at ~110 tok/s on Apple M4 Max, 7.5x llama.cpp, no GPU. From-scratch NEON/SME kernels in Rust.

Rust
#apple-silicon#cpu#gpt-oss#inference#llm#mixture-of-experts#neon#quantization#rust#simd
enhancement help wanted
arizqi/cpubrrr
12

Frontier-class LLM inference on a laptop CPU — gpt-oss:20b at ~110 tok/s on Apple M4 Max, 7.5x llama.cpp, no GPU. From-scratch NEON/SME kernels in Rust.

Rust
#apple-silicon#cpu#gpt-oss#inference#llm#mixture-of-experts#neon#quantization#rust#simd