Qwen proposes GSPO, which optimises policies at the sequence level

Group Sequence Policy Optimization (arXiv v1 2025-07-24, 12 authors) replaces GRPO's token-level importance ratios with ratios based on sequence likelihood and does its clipping, rewarding and…

Date
24 July 2025
Who
Alibaba Qwen team
Confidence
High on what the paper claims (company-reported results)
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): Alibaba Qwen team · Confidence: High on what the paper claims (company-reported results) Group Sequence Policy Optimization (arXiv v1 2025-07-24, 12 authors) replaces GRPO's token-level importance ratios with ratios based on sequence likelihood and does its clipping, rewarding and optimisation at the sequence level; the authors report better training efficiency and performance than GRPO, including more stable RL training of mixture-of-experts models, and a possible simplification of RL infrastructure, and say these properties contributed to the improvements in the latest Qwen3 models (arXiv 2507.18071). It is the production-minded step in the open RL-algorithm line that runs GRPO (DeepSeekMath, 2024-02) to DAPO and Dr. GRPO (2025-03) to CISPO in MiniMax-M1 (2025-06) to GSPO (2025-07), each fixing a bias or instability in the last (token-level clipping, length bias, MoE variance). Which variants frontier labs use in production is mostly undisclosed (Backlog). See also B06 and B08-20. Sources: arXiv 2507.18071

Read it in the deep dive