OpenAI released Deep Research, an o3 model trained with reinforcement learning to browse

OpenAI released a ChatGPT agent that browses and reasons for minutes to write cited reports, powered by an early o3 trained with reinforcement learning on browsing tasks.

Date
2 February 2025
Who
OpenAI
Confidence
High on how it was built (the system card was opened); Medium on the launch date
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Landmark · Significance: 4/5 · Org(s): OpenAI · Confidence: High on how it was built (the system card was opened); Medium on the launch date and benchmark numbers (OpenAI's launch page returned HTTP 403) Primary sources: Deep research system card, 2025-02-25 · OpenAI launch post (403 to the fetcher) · Google, Gemini Deep Research, 2024-12-11 · xAI, Grok 3 and DeepSearch

One-liner. OpenAI released a ChatGPT agent that browses and reasons for minutes to write cited reports, powered by an early o3 trained with reinforcement learning on browsing tasks. It was the first publicly documented RL-trained research agent and linked reasoning models to tool-using agents.

Why it happened. After o1 reasoning models were strong on closed problems but could not look anything up. Google had already shipped Gemini Deep Research on 2024-12-11, a multi-step research plan run on Gemini 1.5 Pro with reports exportable to Google Docs (Google), so the product category existed 53 days earlier, though Google's post does not describe how the model was trained. OpenAI's difference was in the training. On the date, Wikipedia's o3 page gives 2025-02-02 for the launch and its Deep Research page gives 2025-02-03 (citing Reuters of that date), and this book uses 2025-02-02. The launch came 13 days after R1 and two days after o3-mini. I infer that the timing reflects OpenAI's own roadmap for o3 (previewed in 2024-12) and was not a reaction to R1.

The idea. In plain English, the reasoning model itself is trained, with RL, on tasks that require browsing, instead of a general chat model being prompted to call a search tool, so it learns when to search, what to open, how to cross-check and when to stop. How it works (system card): the model is "an early version of OpenAI o3 that is optimized for web browsing." It was trained on new browsing datasets built for research use cases and learned the core browsing skills (searching, clicking, scrolling, interpreting files), how to use a Python tool in a sandbox for calculations and plots, and how to reason through and synthesise many websites, all through reinforcement learning on these browsing tasks. The datasets range from objective tasks with ground-truth answers to open-ended tasks with grading rubrics, and responses were graded against the answers or rubrics using a chain-of-thought model as the grader. It was also trained on o1-era safety data plus browsing-specific safety data, including resistance to malicious instructions found on web pages. In ChatGPT a second, custom-prompted o3-mini summarises the chains of thought, so users see a processed summary, as in B08-16. It launched to Pro users first.

Results (company-reported). Humanity's Last Exam (B08-11a) 26.6% for the o3-based version, as given on Wikipedia's account of OpenAI's announcement (Medium confidence); reports put a run at roughly 5 to 30 minutes.

How it spread. Lag is measured in days from 2025-02-02.

Lab / projectResponseDateLag
GoogleGemini Deep Research on Gemini 1.5 Pro (a research-agent product; its announcement does not describe RL training)2024-12-11-53
xAIDeepSearch launched with Grok 3 (xAI's post dates 2025-02-19)2025-02-17+15
AnthropicClaude 3.7 Sonnet with Claude Code (an agentic coding loop rather than research)2025-02-24+22
OpenAIo3 and o4-mini use every ChatGPT tool inside the chain of thought2025-04-16+73
AnthropicClaude 4 extended thinking with tool use2025-05-22+109
xAIGrok 4, natively trained with tools2025-07-09+157
MoonshotK2 Thinking, 200-300 sequential tool calls2025-11-06+277
DeepSeekV3.2, thinking in tool use2025-12-01+302

Other labs' research-agent products (Perplexity, Anthropic's Research) were not verified here (see the Backlog). A product that users could try, plus a system card that disclosed the recipe in outline (RL on browsing tasks, rubric graders) without the data, carried the idea.

Why it mattered. It is where "search inside the reasoning loop" became a mass-market product, and the clearest early case of RL on a tool-using task that is not maths or code. It narrows the claim, quoted in B08-24, that o3 and o4-mini were the first reasoning models to use tools. That claim is about using all of ChatGPT's tools and images inside the chain of thought, 73 days after this release. It set up the agent stack of B09 and B10.

Nuance, controversy and myths. "Deep research" is a product name several labs use; only OpenAI's system card says the model itself was RL-trained for browsing. The Deep Research model is an early o3 variant, not the o3 that shipped on 2025-04-16 (B08-24). The system card is dated 2025-02-25, three weeks after launch.

Interview kit.

  • 30-second version: Deep Research is an early o3 that OpenAI trained with reinforcement learning on browsing tasks, so it plans searches, reads pages, runs Python and writes a cited report in minutes; it is the first widely used agent built by training the reasoning model itself on tool use.
  • Likely follow-ups: Is it o3 plus a search plugin? → No; the system card describes RL training on browsing datasets with answer or rubric graders. Who was first? → Google shipped a research-agent product on 2024-12-11; OpenAI is the first with a documented RL recipe. Why does it matter for reasoning? → It is the step from reasoning about a prompt to reasoning over the web, which is how reasoning models became agents.
  • Common mistake: Saying it launched with o3 (o3 shipped 73 days later) or that OpenAI invented the deep-research product category.
  • Connect it to: B08-24, B09, B10, B08-08.

Sources. [1] System card, opened and read (pages 1-3). [2] Google's launch post, opened. [3] xAI Grok 3 post, opened. [4] Wikipedia, OpenAI o3 and ChatGPT Deep Research, used as pointers for the launch date and the HLE figure; OpenAI's launch page, Reuters and Business Insider were not fetchable.

Read it in the deep dive