FlashInfer: Kernel Library for LLM Serving

attention cuda distributed-inference gpu jit large-large-models llm-inference moe nvidia pytorch
8 Open Issues Need Help Last updated: Jul 9, 2026

Open Issues Need Help

View All on GitHub
good first issue cute-dsl refactor op: gemm op: comm

FlashInfer: Kernel Library for LLM Serving

Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
good first issue wip op: gemm arch: sm12x

FlashInfer: Kernel Library for LLM Serving

Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
good first issue priority: should have (P1) op: moe

FlashInfer: Kernel Library for LLM Serving

Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
bug good first issue priority: should have (P1) op: moe

FlashInfer: Kernel Library for LLM Serving

Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
help wanted priority: must have (P0) needs-triage op: moe-routing

FlashInfer: Kernel Library for LLM Serving

Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
good first issue op: moe

FlashInfer: Kernel Library for LLM Serving

Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
feature request good first issue op: moe

FlashInfer: Kernel Library for LLM Serving

Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
documentation good first issue

FlashInfer: Kernel Library for LLM Serving

Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch