A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a CPU, never compute the same thing twice: it caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1.

cpu gguf inference kv-cache llama-cpp llm local-llm openai-api self-hosted speculative-decoding
5 Open Issues Need Help Last updated: Jul 27, 2026

Open Issues Need Help

View All on GitHub
help wanted good first issue

A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a CPU, never compute the same thing twice: it caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1.

C++
#cpu#gguf#inference#kv-cache#llama-cpp#llm#local-llm#openai-api#self-hosted#speculative-decoding
help wanted good first issue

A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a CPU, never compute the same thing twice: it caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1.

C++
#cpu#gguf#inference#kv-cache#llama-cpp#llm#local-llm#openai-api#self-hosted#speculative-decoding
enhancement good first issue

A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a CPU, never compute the same thing twice: it caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1.

C++
#cpu#gguf#inference#kv-cache#llama-cpp#llm#local-llm#openai-api#self-hosted#speculative-decoding
good first issue

A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a CPU, never compute the same thing twice: it caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1.

C++
#cpu#gguf#inference#kv-cache#llama-cpp#llm#local-llm#openai-api#self-hosted#speculative-decoding
help wanted good first issue

A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a CPU, never compute the same thing twice: it caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1.

C++
#cpu#gguf#inference#kv-cache#llama-cpp#llm#local-llm#openai-api#self-hosted#speculative-decoding