A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.

0 stars 0 forks 0 watchers Python Apache License 2.0
cuda inference kv-cache llm pytorch vllm
6 Open Issues Need Help Last updated: Aug 23, 2026

Open Issues Need Help

View All on GitHub

A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.

Python
#cuda#inference#kv-cache#llm#pytorch#vllm
help wanted performance

A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.

Python
#cuda#inference#kv-cache#llm#pytorch#vllm
help wanted performance

A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.

Python
#cuda#inference#kv-cache#llm#pytorch#vllm

A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.

Python
#cuda#inference#kv-cache#llm#pytorch#vllm

A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.

Python
#cuda#inference#kv-cache#llm#pytorch#vllm
good first issue

A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.

Python
#cuda#inference#kv-cache#llm#pytorch#vllm