OpenAI released o1-preview and o1-mini, the first public "reasoning model"

OpenAI shipped a model trained with large-scale reinforcement learning to think in a long private chain of thought before answering, and showed that accuracy rises with both training compute and…

Date
12 September 2024
Who
OpenAI
People
Jerry Tworek (listed as "overall" under Leadership in the system card), Noam Brown, Hunter Lightman, Ilge Akkaya, Jakub Pachocki, Mark Chen, Ilya Sutskever (listed as a foundational contributor)
Confidence
High on what was released and claimed; Low on how it was trained (never disclose
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Landmark · Significance: 5/5 · Org(s): OpenAI · People: Jerry Tworek (listed as "overall" under Leadership in the system card), Noam Brown, Hunter Lightman, Ilge Akkaya, Jakub Pachocki, Mark Chen, Ilya Sutskever (listed as a foundational contributor) · Confidence: High on what was released and claimed; Low on how it was trained (never disclosed) Primary sources: Learning to reason with LLMs (blog, 2024-09-12) · o1 System Card, arXiv 2412.16720 · Simon Willison's launch notes

One-liner. OpenAI shipped a model trained with large-scale reinforcement learning to think in a long private chain of thought before answering, and showed that accuracy rises with both training compute and thinking time.

Why it happened. o1 grew out of a research bet that predates the public "pretraining is slowing" story. The slowdown narrative arrived afterwards. The Information reported on 2024-11-09 that OpenAI's next flagship, Orion, showed smaller gains than the jump from GPT-3 to GPT-4 and that OpenAI was turning to synthetic data and heavier post-training (TechCrunch, 2024-11-09; background in B03 and B04). I infer that the slowdown story became the public rationale for a reasoning bet that was already underway. Noam Brown, a game-AI researcher known for the poker and diplomacy systems Libratus, Pluribus and Cicero, argued that letting a model search or "think" for longer at inference could be worth orders of magnitude of model scale, as it had been in games. In a Latent Space interview Brown dates the first clear signs that the reasoning paradigm worked at OpenAI to around October-November 2023, and in a Sequoia Training Data conversation Brown and Hunter Lightman describe many attempts that failed before a version in which the model began to backtrack and correct its own errors without being taught to, which Lightman describes as the point at which the work began to look like it would succeed. The earlier public footprints of this work (the Q\* leak of 2023-11, the "Strawberry" report of 2024-07, process supervision in "Let's Verify Step by Step") are in B07. The system card's authorship page lists Ilya Sutskever, Hyung Won Chung, Jason Wei, Karl Cobbe and Vineet Kosaraju among the foundational contributors, with Jakub Pachocki, Jerry Tworek (overall), Mark Chen and Lukasz Kaiser among the leadership. I infer that the team mixed RLHF, process-supervision and scaling expertise.

The idea. Instead of answering in one pass, the model first writes a long internal chain of thought (CoT) that the user does not see, then answers. It is trained by reinforcement learning (RL) so that useful thinking, such as decomposing a problem, noticing a mistake and trying a different route, is reinforced because it leads to correct answers. How it works: OpenAI disclosed almost nothing about the algorithm, reward design, data or compute. What it did state was that a large-scale RL algorithm trains the CoT, that performance improves consistently with more RL compute at training time and more thinking time at test time, and that the constraints on scaling this differ substantially from LLM pretraining. The blog's headline plots put compute on a log axis. The API exposed the cost directly. Hidden "reasoning tokens" are billed as output tokens, OpenAI suggested budgeting about 25,000 of them for hard prompts, and output caps rose to 32,768 tokens for o1-preview and 65,536 for o1-mini (per Willison). Pricing was $15/$60 per million input/output tokens for o1-preview and $3/$12 for o1-mini, versus $5/$15 for GPT-4o as The Batch quoted it (The Batch); that is GPT-4o's original launch price, and later GPT-4o listings are given as $2.50 / $10 (Wikipedia's GPT-4.5 article, a pointer; OpenAI's pricing archive was not retrievable), which is the comparator B08-36 implies.

Results. All figures are company-reported, for the then-unreleased full o1 model, not o1-preview. On the 2024 AIME (a 15-question US high-school olympiad qualifier), GPT-4o averaged about 12%; o1 averaged 74% with one sample per problem, 83% with consensus over 64 samples, and 93% when re-ranking 1,000 samples with a learned scorer. It reached the 89th percentile on Codeforces and was reported above human PhD-level accuracy on GPQA Diamond (hard graduate-level science multiple-choice questions; about 78% versus roughly 70% for the human experts OpenAI recruited).

The hidden chain of thought. OpenAI chose not to show the raw CoT, citing user experience, competitive advantage and the option to monitor the CoT for manipulation, which in turn meant not training policy compliance into the CoT itself. Users saw a model-written summary instead, and API users saw nothing but the token count. OpenAI also enforced the policy. Users who probed o1 for its reasoning reportedly received warnings that violations could cost them access (Ars Technica, 2024-09-16, as summarised on Wikipedia; the Ars page was not fetchable). The o1 system card's "CoT deception monitoring" ran a GPT-4o monitor over 102,443 o1-preview chains of thought and flagged 0.17% as deceptive, mostly hallucinated policy refusals and made-up references (system card §4.3). The competitive-advantage reason is the one that mattered later, because it is the policy DeepSeek reversed in practice and OpenAI partially walked back (B08-16).

How it spread. Lag is measured in days from 2024-09-12.

Lab / projectResponseDateLag
Open-O1 community projectOpen-source o1-style long-CoT fine-tuning (SFT data for "CoT activation"); initial release 2024-10-05, OpenO1-Qwen-7B-v0.1 on 2024-10-09 (GitHub)2024-10-05+23
GAIR group (Pengfei Liu et al.)"O1 Replication Journey, Part 1", which proposed "journey learning" from search trajectories and tried early distillation2024-10-08 (arXiv 2410.18982)+26
DeepSeekR1-Lite-Preview, visible CoT, claimed o1-preview level on AIME/MATH2024-11-20+69
Alibaba MarcoPoloMarco-o1 (CoT fine-tuning plus MCTS, aimed at open-ended problems; arXiv 2411.14405)2024-11-21+70
Alibaba QwenQwQ-32B-Preview2024-11-28+77
GoogleGemini 2.0 Flash Thinking Experimental2024-12-19+98
DeepSeek / MoonshotR1, Kimi k1.5 with full recipes2025-01-20+130
xAIGrok 3 (Think)2025-02-17+158
AnthropicClaude 3.7 Sonnet extended thinking2025-02-24+165
GoogleGemini 2.5 Pro, thinking by default2025-03-25+194
MistralMagistral2025-06-10+271
MetaMobileLLM-R1 (small, SFT-only reasoning models)2025-09-12+365
MetaMuse Spark, first flagship reasoning model2026-04-08+573

Two things accelerated it. One was the existence proof (a capability, once shown possible, is easier to chase), and the other was that OpenAI exposed the behaviour at the API, so rivals could study outputs and, reportedly, distil them (B08-14). The missing recipe slowed it. Nobody outside OpenAI had a working RL-on-CoT pipeline until DeepSeek published one.

Why it mattered. It added inference-time compute as a second way to scale. Quality became a function of how much you are willing to spend per query, which created the effort and budget settings of B08-06, the multi-dollar-per-task evaluation regimes of B08-08, and the capital argument for more inference hardware (B17, B18). It also made post-training RL on verifiable tasks the main lever, which fed agents (B10).

Nuance, controversy and myths. "o1 was GPT-4 with a prompt" is wrong, because it is a separately RL-trained model. "OpenAI invented chain-of-thought reasoning" is also wrong, since CoT prompting is Google 2022 and the o1 contribution is RL at scale. OpenAI's raw CoT was never released, so every claim about o1's internals (search, Monte Carlo tree search (MCTS), process reward models) is outside inference; DeepSeek's later findings that plain outcome-reward RL suffices undercut the more elaborate guesses.

Interview kit.

  • 30-second version: o1 is a language model trained with reinforcement learning to produce a long hidden chain of thought, so accuracy on hard math, code and science problems goes up the more compute you spend at inference as well as at training.
  • Likely follow-ups: Why hide the thinking? → Safety monitoring (do not train the CoT to look good) and competitive distillation risk. What is a "reasoning token"? → A billed output token the user never sees. Was it search? → Unknown; OpenAI never said, and R1 showed search is not required.
  • Common mistake: Saying o1 was released with o3's numbers; the famous AIME/Codeforces results are for the full o1, not o1-preview.
  • Connect it to: B07, B05, B08-08, B22.

Sources. [1] OpenAI blog above (OpenAI-hosted pages returned HTTP 403 to the research fetcher; figures cross-checked via a search extract of the page and The Batch). [2] System card, opened and read locally. [3] Willison notes. [4] Latent Space, Noam Brown. [5] Sequoia Training Data. [6] TechCrunch on the Orion slowdown report (opened). [7] Open-O1 repository (opened).

Read it in the deep dive