sipllm — a dependency-free streaming GGUF LLM inference engine in C++17. Runs quantized Llama models on-device with flat RAM (it sips weights off disk), Ollama-style CLI, validated layer-by-layer against llama.cpp.

arm64 cpp cpp17 edge-ai gguf inference llama llama-cpp llm machine-learning neon on-device quantization tinyllama transformer
8 Open Issues Need Help Last updated: Jul 23, 2026

Open Issues Need Help

View All on GitHub
enhancement good first issue architecture

sipllm — a dependency-free streaming GGUF LLM inference engine in C++17. Runs quantized Llama models on-device with flat RAM (it sips weights off disk), Ollama-style CLI, validated layer-by-layer against llama.cpp.

C++
#arm64#cpp#cpp17#edge-ai#gguf#inference#llama#llama-cpp#llm#machine-learning#neon#on-device#quantization#tinyllama#transformer
enhancement good first issue architecture

sipllm — a dependency-free streaming GGUF LLM inference engine in C++17. Runs quantized Llama models on-device with flat RAM (it sips weights off disk), Ollama-style CLI, validated layer-by-layer against llama.cpp.

C++
#arm64#cpp#cpp17#edge-ai#gguf#inference#llama#llama-cpp#llm#machine-learning#neon#on-device#quantization#tinyllama#transformer

sipllm — a dependency-free streaming GGUF LLM inference engine in C++17. Runs quantized Llama models on-device with flat RAM (it sips weights off disk), Ollama-style CLI, validated layer-by-layer against llama.cpp.

C++
#arm64#cpp#cpp17#edge-ai#gguf#inference#llama#llama-cpp#llm#machine-learning#neon#on-device#quantization#tinyllama#transformer

sipllm — a dependency-free streaming GGUF LLM inference engine in C++17. Runs quantized Llama models on-device with flat RAM (it sips weights off disk), Ollama-style CLI, validated layer-by-layer against llama.cpp.

C++
#arm64#cpp#cpp17#edge-ai#gguf#inference#llama#llama-cpp#llm#machine-learning#neon#on-device#quantization#tinyllama#transformer

sipllm — a dependency-free streaming GGUF LLM inference engine in C++17. Runs quantized Llama models on-device with flat RAM (it sips weights off disk), Ollama-style CLI, validated layer-by-layer against llama.cpp.

C++
#arm64#cpp#cpp17#edge-ai#gguf#inference#llama#llama-cpp#llm#machine-learning#neon#on-device#quantization#tinyllama#transformer

sipllm — a dependency-free streaming GGUF LLM inference engine in C++17. Runs quantized Llama models on-device with flat RAM (it sips weights off disk), Ollama-style CLI, validated layer-by-layer against llama.cpp.

C++
#arm64#cpp#cpp17#edge-ai#gguf#inference#llama#llama-cpp#llm#machine-learning#neon#on-device#quantization#tinyllama#transformer
good first issue

sipllm — a dependency-free streaming GGUF LLM inference engine in C++17. Runs quantized Llama models on-device with flat RAM (it sips weights off disk), Ollama-style CLI, validated layer-by-layer against llama.cpp.

C++
#arm64#cpp#cpp17#edge-ai#gguf#inference#llama#llama-cpp#llm#machine-learning#neon#on-device#quantization#tinyllama#transformer
good first issue

sipllm — a dependency-free streaming GGUF LLM inference engine in C++17. Runs quantized Llama models on-device with flat RAM (it sips weights off disk), Ollama-style CLI, validated layer-by-layer against llama.cpp.

C++
#arm64#cpp#cpp17#edge-ai#gguf#inference#llama#llama-cpp#llm#machine-learning#neon#on-device#quantization#tinyllama#transformer