Qwen proposes GSPO, which optimises policies at the sequence level
Group Sequence Policy Optimization (arXiv v1 2025-07-24, 12 authors) replaces GRPO's token-level importance ratios with ratios based on sequence likelihood and does its clipping, rewarding and…
- Date
- 24 July 2025
- Who
- Alibaba Qwen team
- Confidence
- High on what the paper claims (company-reported results)
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): Alibaba Qwen team · Confidence: High on what the paper claims (company-reported results) Group Sequence Policy Optimization (arXiv v1 2025-07-24, 12 authors) replaces GRPO's token-level importance ratios with ratios based on sequence likelihood and does its clipping, rewarding and optimisation at the sequence level; the authors report better training efficiency and performance than GRPO, including more stable RL training of mixture-of-experts models, and a possible simplification of RL infrastructure, and say these properties contributed to the improvements in the latest Qwen3 models (arXiv 2507.18071). It is the production-minded step in the open RL-algorithm line that runs GRPO (DeepSeekMath, 2024-02) to DAPO and Dr. GRPO (2025-03) to CISPO in MiniMax-M1 (2025-06) to GSPO (2025-07), each fixing a bias or instability in the last (token-level clipping, length bias, MoE variance). Which variants frontier labs use in production is mostly undisclosed (Backlog). See also B06 and B08-20. Sources: arXiv 2507.18071