Karpathy's "State of GPT"

At Microsoft Build 2023 (session BRK216HFS) Andrej Karpathy gave a widely shared practitioner walk-through of how a GPT-style assistant is built (view counts not checked).

Date
25 May 2023
Who
OpenAI (speaker), Microsoft Build (venue)
People
Andrej Karpathy
Confidence
Medium (the session's existence, topic and recording date are confirmed; I could
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): OpenAI (speaker), Microsoft Build (venue) · People: Andrej Karpathy · Confidence: Medium (the session's existence, topic and recording date are confirmed; I could not open the slides or transcribe the talk, so I do not attribute specific claims to it) At Microsoft Build 2023 (session BRK216HFS) Andrej Karpathy gave a widely shared practitioner walk-through of how a GPT-style assistant is built (view counts not checked). It covered tokenization and pretraining, then supervised fine-tuning, reward modeling and reinforcement learning, followed by practical prompting and tool-use advice (session listing, a mirror of the Build page; the recording was posted to YouTube on 2023-05-25, video; slides are linked from Karpathy's site). It belongs here because its four-stage picture (pretraining, SFT, reward modeling, RL) is a common way to explain where RLHF sits. It is an explanatory talk and not a research result, so a reader should check any specific claim against the slides before attributing it (Backlog). It sits between Schulman's talk (B05-32) and Karpathy's 2024 "vibe check" remark on reward models (B05-18). Sources: Build session page · YouTube recording

Read it in the deep dive