Open Issues Need Help
View All on GitHub Llama-lineage CPU operations — RmsNorm, Swiglu, Rope, TokenEmbedding about 2 hours ago
enhancement good first issue
A C++23 module-based DNN library for GPU-first LLM inference — explicit forward passes, no hidden execution engine, work at the metal. Validated token-for-token on Gemma 4 Unified, Llama 3.x and GPT-2 with compile-time FP8/FP4 weight quantization.
C++
#cpp20-modules#cpp23#cuda#deep-learning#flashattention#fp4-quantization#fp8-quantization#gemma4-12b#gpt-2#gpu#inference#inference-server#llama3-2#llm#local-ai#mnist#neural-network#tensors#training#transformers
enhancement good first issue
A C++23 module-based DNN library for GPU-first LLM inference — explicit forward passes, no hidden execution engine, work at the metal. Validated token-for-token on Gemma 4 Unified, Llama 3.x and GPT-2 with compile-time FP8/FP4 weight quantization.
C++
#cpp20-modules#cpp23#cuda#deep-learning#flashattention#fp4-quantization#fp8-quantization#gemma4-12b#gpt-2#gpu#inference#inference-server#llama3-2#llm#local-ai#mnist#neural-network#tensors#training#transformers