Forge-researchlab

Forge-researchlab/forge-kernels

Forge kernels — fused Triton kernels for faster, leaner LLM fine-tuning. One-call patching into Hugging Face models via forge.patch(model), with FSDP2 multi-GPU support and a growing set of kernels and architectures. Every speedup backed by a committed benchmark. Apache-2.0.

1 good first / help-wanted issue · Python · last activity Aug 18, 2026

9 stars 4 forks 9 watchers Python Apache License 2.0
cuda fine-tuning fsdp gemma gpu-kernels gpu-optimization huggingface llm lora pytorch qwen transformers triton
1 Open Issue Need Help Last updated: Aug 18, 2026

Open Issues Need Help

View All on GitHub
enhancement good first issue
Forge-researchlab/forge-kernels
9

Forge kernels — fused Triton kernels for faster, leaner LLM fine-tuning. One-call patching into Hugging Face models via forge.patch(model), with FSDP2 multi-GPU support and a growing set of kernels and architectures. Every speedup backed by a committed benchmark. Apache-2.0.

Python
#cuda#fine-tuning#fsdp#gemma#gpu-kernels#gpu-optimization#huggingface#llm#lora#pytorch#qwen#transformers#triton