OpenAI publishes Proximal Policy Optimization, which became the default RLHF optimizer

PPO is a simple policy-gradient method that takes several gradient steps per batch of data while limiting how far the policy moves; it became the default optimizer for RLHF from 2019 to roughly 2023.

Date
20 July 2017
Who
OpenAI
People
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Landmark · Significance: 4/5 · Org(s): OpenAI · People: John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov · Confidence: High Primary sources: arXiv:1707.06347 (v1 2017-07-20)

One-liner. PPO is a simple policy-gradient method that takes several gradient steps per batch of data while limiting how far the policy moves; it became the default optimizer for RLHF from 2019 to roughly 2023.

Why it happened. Schulman's earlier TRPO limited how far each update could move the policy but was complicated to implement. PPO was meant to keep most of the benefit with first-order code (paper). It was a general RL algorithm, tested on simulated robotics and Atari; it was not designed for language or for human feedback.

The idea. Collect experience, then optimize a "surrogate" objective with stochastic gradient ascent for multiple epochs on the same batch, while clipping the ratio between new and old action probabilities so a single update cannot move the policy too far.

How it works (in RLHF). The language model is the policy; each prompt is a one-step "bandit" episode; the reward-model score is the reward; a per-token KL penalty to the supervised starting model discourages drift; a value function (initialized from the reward model in InstructGPT) reduces variance (InstructGPT §3.5). In practice at least a policy, a frozen reference model, a reward model and a value model must be held in memory, which is part of why open-source RLHF stayed hard until libraries appeared (B05-17) and why DPO was attractive (B05-35). Cost relative to pretraining was small, at 60 petaflop/s-days for the 175B InstructGPT PPO-ptx run versus 3,640 for GPT-3 pretraining (InstructGPT §5.1).

How it spread. First used for language with human feedback at OpenAI in 2019 (Ziegler, B05-05), then Stiennon, InstructGPT, ChatGPT (OpenAI's post says PPO; blog), Llama 2-Chat (PPO applied on top of rejection sampling, but only in the last of its five RLHF versions; B05-37). Open-source RLHF libraries (TRL, trlX, DeepSpeed-Chat) were PPO implementations first. Replacements for the alignment use (DPO and relatives) are in B06; reasoning-style RL moved to other variants (B07, B08).

Why it mattered. It gave the OpenAI team a stable, already-debugged optimizer that could be pointed at a reward model. Choosing PPO was plausibly an organizational convenience (its author, Schulman, was an OpenAI co-founder working on exactly this) as much as a technical necessity. Inference: the later move to DPO and verifier-based RL shows that PPO was one workable choice and that "RLHF" does not require it.

Nuance, controversy and myths. Christiano et al. (2017) used A2C and TRPO, not PPO (B05-02). "RLHF = PPO" is a common conflation; the term describes the data and reward-model loop, and OpenAI's own text-davinci-002 used human data but a different method (B05-19).

Interview kit.

  • 30-second version: PPO is the RL algorithm Schulman's team published in 2017; it limits how much the policy changes per update, and it was used to optimize language models against a reward model in InstructGPT and ChatGPT.
  • Likely follow-ups: Why a KL penalty? → Without it the policy exploits the reward model and drifts into gibberish that scores well (B05-18). Is PPO still used? → Variants and alternatives dominate newer pipelines; see B06.
  • Common mistake: Saying the 2017 human-preferences paper used PPO.
  • Connect it to: B05-02, B05-13, B05-35.

Sources. 1. Schulman et al., arXiv:1707.06347 · 2. Ouyang et al., arXiv:2203.02155 · 3. OpenAI, "Introducing ChatGPT", 2022-11-30

Read it in the deep dive