LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

amd cuda fast inference kv-cache llm pytorch rocm speed vllm
63 Open Issues Need Help Last updated: Jul 8, 2026

Open Issues Need Help

View All on GitHub

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
documentation good first issue help wanted onboarding-2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted stale RFC

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
documentation good first issue help wanted

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted onboarding-2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted onboarding-2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted onboarding-2026 area/lint

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
enhancement help wanted discussion stale RFC optimization

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
help wanted Testing stale backend

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted stale new feature

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
enhancement help wanted Testing stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
help wanted stale new feature hardware

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
enhancement good first issue help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
documentation good first issue stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
bug good first issue help wanted

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
bug good first issue help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
help wanted discussion stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted stale controller

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted backend

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted controller

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
enhancement help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
enhancement help wanted stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue help wanted Testing stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
enhancement good first issue Refactoring stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm
good first issue stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

AI Summary: The `save_decode_cache` configuration variable, which is intended to control whether the decode KV cache is saved, is currently not effective in `lmcache/vllm v1`. The task involves modifying `lmcache/integration/vllm/vllm_v1_adapter.py` to properly implement this optional saving functionality based on the config.

Complexity: 1/5
good first issue stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

AI Summary: This GitHub issue describes a bug in LMCache where, even after successfully loading KV cache for hit tokens, it redundantly attempts to store those same tokens again. This leads to unnecessary write operations to the cache, as observed when sending identical requests and seeing logs indicating all tokens are being stored instead of just the newly generated ones.

Complexity: 2/5
bug good first issue stale

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

AI Summary: Debug a peer-to-peer (P2P) key-value (KV) cache sharing issue in LMCache, an LLM serving engine. The problem is that a second instance of the engine cannot load caches stored on the disk of the first instance, even though the first instance successfully stored and retrieved the caches. The task involves analyzing the provided Python code and debugging the P2P mechanism to ensure proper cache loading across instances.

Complexity: 4/5
bug good first issue help wanted

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

AI Summary: Optimize the CI/CD pipeline for the LMCache project. This involves improving the handling of unit test failures (stopping the pipeline on failure), configuring appropriate concurrency for unit tests (single GPU) and end-to-end tests (two GPUs), and potentially correcting concurrency settings for unit and integration tests.

Complexity: 4/5
help wanted

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

AI Summary: Write documentation for AMD GPU support in LMCache, mirroring the existing vLLM AMD ROCm documentation. This involves detailing how to install and utilize AMD GPUs with LMCache.

Complexity: 3/5
documentation help wanted

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

AI Summary: Add a configuration parameter to LMCache to allow users to specify the Infinistore link type (Ethernet or Infiniband) at runtime, instead of hardcoding it to Ethernet. This will improve flexibility and allow users to leverage Infiniband connections if available.

Complexity: 3/5
good first issue

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm

AI Summary: Update the configuration files (lmcache-prefiller-config.yaml and lmcache-decoder-config.yaml) within the LMCache project's examples (1p1d and xp1d) to replace instances of `nixl_peer_host` and `nixl_peer_port` with `nixl_receiver_host` and `nixl_receiver_port`, respectively. This involves correcting the naming convention for the host and port used in the disagg_prefill configuration.

Complexity: 2/5
good first issue

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python
#amd#cuda#fast#inference#kv-cache#llm#pytorch#rocm#speed#vllm