A C++23 module-based DNN library for GPU-first LLM inference — explicit forward passes, no hidden execution engine, work at the metal. Validated token-for-token on Gemma 4 Unified, Llama 3.x and GPT-2 with compile-time FP8/FP4 weight quantization.

cpp20-modules cpp23 cuda deep-learning flashattention fp4-quantization fp8-quantization gemma4-12b gpt-2 gpu inference inference-server llama3-2 llm local-ai mnist neural-network tensors training transformers
2 Open Issues Need Help Last updated: Jul 21, 2026

Open Issues Need Help

View All on GitHub

A C++23 module-based DNN library for GPU-first LLM inference — explicit forward passes, no hidden execution engine, work at the metal. Validated token-for-token on Gemma 4 Unified, Llama 3.x and GPT-2 with compile-time FP8/FP4 weight quantization.

C++
#cpp20-modules#cpp23#cuda#deep-learning#flashattention#fp4-quantization#fp8-quantization#gemma4-12b#gpt-2#gpu#inference#inference-server#llama3-2#llm#local-ai#mnist#neural-network#tensors#training#transformers

A C++23 module-based DNN library for GPU-first LLM inference — explicit forward passes, no hidden execution engine, work at the metal. Validated token-for-token on Gemma 4 Unified, Llama 3.x and GPT-2 with compile-time FP8/FP4 weight quantization.

C++
#cpp20-modules#cpp23#cuda#deep-learning#flashattention#fp4-quantization#fp8-quantization#gemma4-12b#gpt-2#gpu#inference#inference-server#llama3-2#llm#local-ai#mnist#neural-network#tensors#training#transformers