Open Issues Need Help
View All on GitHub Next fix for gemm-allreduce two-shot 3 days ago
good first issue cute-dsl refactor op: gemm op: comm
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
DGX Spark (SM121) Current Support Audit 5 days ago
good first issue wip op: gemm arch: sm12x
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
Potentially superfluous check that disables non gated activations in the cutlass fused moe API 3 months ago
good first issue priority: should have (P1) op: moe
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
MoE autotune print a lot failed kernel on SM120 3 months ago
bug good first issue priority: should have (P1) op: moe
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
help wanted priority: must have (P0) needs-triage op: moe-routing
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
BF16 hidden_states for trtllm_fp4_block_scale_moe 4 months ago
good first issue op: moe
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
Autotuning failsafe fallback to top1 tactic 6 months ago
feature request good first issue op: moe
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch
Inaccurate API Docstrings for Attention Prefill 8 months ago
documentation good first issue
FlashInfer: Kernel Library for LLM Serving
Python
#attention#cuda#distributed-inference#gpu#jit#large-large-models#llm-inference#moe#nvidia#pytorch