The shoggoth meme and the "just a mask" debate
One month after ChatGPT, the Twitter user @TetraspaceWest drew two Lovecraftian shoggoths, "GPT-3" and "GPT-3 + RLHF", the second holding a tiny smiley-face mask on a tentacle (Lambert's RLHF book…
- Date
- 30 December 2022
- Who
- community; Janus, Anthropic and OpenAI researchers
- Confidence
- Medium (meme provenance); High (quotes and dates of the technical positions)
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): community; Janus, Anthropic and OpenAI researchers · Confidence: Medium (meme provenance); High (quotes and dates of the technical positions) One month after ChatGPT, the Twitter user @TetraspaceWest drew two Lovecraftian shoggoths, "GPT-3" and "GPT-3 + RLHF", the second holding a tiny smiley-face mask on a tentacle (Lambert's RLHF book notes; Know Your Meme); the New York Times columnist Kevin Roose wrote about it in May 2023 (I know this only from search-result summaries; the article was not opened). The meme claims that RLHF changes the mask the model wears and leaves what is underneath unchanged. Positions. Mask-like evidence: Janus's "Simulators" framing treats a base model as a simulator of many characters that post-training narrows to one, the helpful-honest-harmless assistant; LIMA's "superficial alignment hypothesis" says alignment mostly teaches format and style (B05-33); jailbreak papers show that safety training can be undone cheaply, since 10 adversarial fine-tuning examples costing under $0.20 broke GPT-3.5 Turbo's guardrails and automatically found adversarial suffixes transfer across aligned models (Qi et al., 2023-10-05; Zou et al., 2023-07). Not-just-a-mask evidence: RLHF measurably changes behavior distributions and creates failure modes (sycophancy, stated preferences) that look like learned dispositions (B05-22, B05-39); InstructGPT showed a 1.3B RLHF model beating a 175B base on user preference (B05-13). Insiders do not settle it. The summary line on host Dwarkesh Patel's 2024-05-15 interview with John Schulman says "how posttraining tames the shoggoth", but that is the host's framing (the episode title is "Reasoning, RLHF, & plan for 2027 AGI", and I did not see Schulman use the term in the transcript text I read) (Dwarkesh, 2024-05-15); Christiano described RLHF as a first step, neither a mask nor a solution (Dwarkesh, 2023-10-31). Inference: the meme is rhetorically useful but under-specified; the empirical question is how deep the behavior change goes under distribution shift, which is B21 and B22 territory. Sources: Lambert · KYM · Qi et al. · Zou et al. · Dwarkesh/Schulman