Alibaba released the open-weight reasoning models QwQ-32B-Preview and Marco-o1

QwQ-32B-Preview was an early and widely used open-weight long-CoT model (Marco-o1 and others were smaller or earlier).

Date
28 November 2024
Who
Alibaba (Qwen team; MarcoPolo team)
Confidence
High
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): Alibaba (Qwen team; MarcoPolo team) · Confidence: High QwQ-32B-Preview was an early and widely used open-weight long-CoT model (Marco-o1 and others were smaller or earlier). It had 32B parameters and reported 65.2% GPQA, 50.0% AIME, 90.6% MATH-500 and 50.0% LiveCodeBench, with the team candidly listing language mixing, endless recursive loops and weak common-sense behaviour as limits (Qwen blog, 2024-11-28). It landed 77 days after o1-preview (Alibaba's separate MarcoPolo group had posted the smaller Marco-o1, built on CoT fine-tuning plus MCTS, a week earlier, as described in arXiv 2411.14405). QwQ's outputs became a favourite teacher. Berkeley's Sky-T1 distilled 17K QwQ traces (B08-13), and the R1 paper uses QwQ-32B-Preview as its open baseline. Qwen's RL-trained successor is QwQ-32B (2025-03-06).

Read it in the deep dive