Alibaba released the open-weight reasoning models QwQ-32B-Preview and Marco-o1
QwQ-32B-Preview was an early and widely used open-weight long-CoT model (Marco-o1 and others were smaller or earlier).
- Date
- 28 November 2024
- Who
- Alibaba (Qwen team; MarcoPolo team)
- Confidence
- High
- Deep dive
- Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them
Tier: Supporting · Significance: 3/5 · Org(s): Alibaba (Qwen team; MarcoPolo team) · Confidence: High QwQ-32B-Preview was an early and widely used open-weight long-CoT model (Marco-o1 and others were smaller or earlier). It had 32B parameters and reported 65.2% GPQA, 50.0% AIME, 90.6% MATH-500 and 50.0% LiveCodeBench, with the team candidly listing language mixing, endless recursive loops and weak common-sense behaviour as limits (Qwen blog, 2024-11-28). It landed 77 days after o1-preview (Alibaba's separate MarcoPolo group had posted the smaller Marco-o1, built on CoT fine-tuning plus MCTS, a week earlier, as described in arXiv 2411.14405). QwQ's outputs became a favourite teacher. Berkeley's Sky-T1 distilled 17K QwQ traces (B08-13), and the R1 paper uses QwQ-32B-Preview as its open baseline. Qwen's RL-trained successor is QwQ-32B (2025-03-06).