Reproducing and studying RL algorithms for LLM agents, including PPO, GRPO, GSPO, DAPO, OPD and beyond.

89 stars 8 forks 89 watchers Python Apache License 2.0
1 Open Issue Need Help Last updated: Jul 30, 2026

Open Issues Need Help

View All on GitHub
help wanted

Reproducing and studying RL algorithms for LLM agents, including PPO, GRPO, GSPO, DAPO, OPD and beyond.

Python