QwQ-32B, Qwen3 and the hybrid-thinking reversal

Alibaba's QwQ-32B (blog 2025-03-06, Apache 2.0) claimed performance comparable to DeepSeek-R1 (671B total, 37B active) at 32B parameters, trained from a cold-start checkpoint with outcome-based RL…

Date
6 March 2025
Who
Alibaba Qwen
Confidence
High for blog claims; Medium for the reason for the reversal
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): Alibaba Qwen · Confidence: High for blog claims; Medium for the reason for the reversal Alibaba's QwQ-32B (blog 2025-03-06, Apache 2.0) claimed performance comparable to DeepSeek-R1 (671B total, 37B active) at 32B parameters, trained from a cold-start checkpoint with outcome-based RL (an accuracy verifier for maths and a code-execution server), then a general-reward stage (Qwen blog). On 2025-04-29 Qwen3 followed (+99 days after R1) as open weights under Apache 2.0, from a 235B-A22B flagship to dense 0.6B-32B models, trained on about 36T tokens in four stages (long-CoT cold start, reasoning RL, "thinking mode fusion," general RL) so a single model could switch between thinking and non-thinking with /think and /no_think and a thinking budget (Qwen3 blog). Then in July 2025 the team announced it would stop using hybrid mode and train separate Instruct and Thinking models (the "2507" releases, with 256K context), saying that better quality mattered more than unification at that point, while continuing hybrid research (The Register, 2025-07-31; the Instruct-2507 model card confirms it is non-thinking only), a rare public reversal that implies fusing the two behaviours in one open model cost quality. The line continued in 2026 with Qwen3.5 (2026-02), proprietary 3.7 releases (2026-05 and 06), then Qwen3.8-Max (general availability 2026-08-03) and an open-weights sibling, Qwen3.8-2.4T-A95B (2026-08-12), a text-only 2.4T-parameter model whose model card says every interaction requires thinking mode, with no non-thinking mode (model card; dates via Wikipedia's Qwen article, a pointer). In order, that is hybrid (Qwen3), separate Instruct and Thinking models (2025-07), then a thinking-only open flagship (2026-08). The RL algorithm behind Qwen3 is in B08-34a and the wider 2026 line in B08-43c. See also B08-06, B19, B20.

Read it in the deep dive