Open Issues Need Help
View All on GitHubA vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.
A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.
A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.
A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.
A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.
A vLLM-style LLM inference engine written from scratch on a 4 GB laptop GPU: paged KV cache, continuous batching, prefix caching, preemption, CUDA graphs, and a streaming server. 3.65x HF generate() on realistic workloads.