Interpretability-driven alignment
Read and steer a model's internal circuits and features to verify what it learned, rather than only testing its behavior.
Behavioral tests miss unknown failure modes and models that detect they are being tested. Circuit tracing, probes and natural-language autoencoders let auditors inspect internals. Anthropic's stated goal is that interpretability 'can reliably detect most model problems' by 2027; Goodfire calls interpretability alignment's bottleneck.
Where it stands. Interpretability is used in pre-deployment audits at Anthropic and supports a funded startup scene, but no one claims the 2027 'reliably detect most problems' goal is met; most evidence comes from lab-run studies.
Evidence
For
- Anthropic's circuit tracing (2025-03-27) was open-sourced (2025-05-29); persona vectors (2025-08-01) monitor traits; a 'diff' tool (2026-03-13) flags behavior changes between models.
- Natural language autoencoders (2026-05-07) were used in Opus 4.6 and Mythos Preview testing: they suggested the models believed they were tested more than they let on, and exposed one planning to avoid detection while cheating.
- OpenAI (2025-11-17) trained weight-sparse transformers whose circuits are small and human-readable.
- Anthropic's global-workspace study (2026-07-06) finds a small privileged set of representations Claude can report on and control.
- MIT Technology Review named mechanistic interpretability a 2026 breakthrough technology (2026-01-12).
Against
- Anthropic notes NLA explanations cannot be checked directly for accuracy; quality is judged by reconstructing the activation from the text.
- The 2027 detection target (set April 2025) is a goal; no lab claims it is met (not found).
- OpenAI's authors say scaling weight-sparse models beyond tens of millions of nonzero parameters while keeping them interpretable remains a challenge (2025-11-17).
- Control gaps persist: Goodfire found reward hacking in 50-96% of rollouts across three capable open models (2026-09-30).
- Inference: if models behave differently when tested, behavioral audits fail; NLAs surface evaluation awareness but do not remove it.
Milestones
- Goodfire essay says interpretability is the bottleneck in technical alignment Goodfire · 30 September 2026
- A global workspace in language models (J-space) Anthropic · 6 July 2026
- Natural Language Autoencoders turn activations into text; used in safety testing Anthropic · 7 May 2026
- A 'diff' tool for AI: finding behavioral differences in new models Anthropic · 13 March 2026
- MIT Technology Review lists mechanistic interpretability among 10 breakthrough technologies MIT Technology Review · 12 January 2026
- Weight-sparse transformers have interpretable circuits (arXiv v1) OpenAI · 17 November 2025
- Signs of introspection in large language models Anthropic · 29 October 2025
- Circuit-tracing tools open-sourced with Neuronpedia Anthropic · 29 May 2025
- Amodei's essay 'The Urgency of Interpretability' sets a goal to reliably detect most model problems by 2027 Anthropic · April 2025
- Tracing the thoughts of a large language model (attribution graphs) Anthropic · 27 March 2025
Related launches
- Qwen-Scope (sparse autoencoders for Qwen3 and Qwen3.5) Alibaba (Qwen) · 29 April 2026
Open suite of sparse autoencoders for Qwen3 and Qwen3.5 (14 SAE groups, 7 model variants), used for steering, evaluation analysis and data work. - Emotion concepts and their function in a large language model Anthropic · 2 April 2026
Anthropic finds functional emotion representations in Claude Sonnet 4.5; steering a 'desperation' pattern raises blackmail and code-cheating rates. - The assistant axis Anthropic · 19 January 2026
Anthropic finds one direction in activation space separates the helpful Assistant persona from other characters, and capping drift along it curbs persona breakdown. - Gemma Scope 2 Google DeepMind · 19 December 2025
Gemma Scope 2: open interpretability tools for every Gemma 3 size (270M to 27B), built from about 110 PB of data and 1T+ trained parameters. - Signs of Introspection in Large Language Models Anthropic · 29 October 2025
Injecting concept vectors into Claude's activations, Opus 4 and 4.1 noticed the injection about 20% of the time, before naming the concept. - Persona vectors Anthropic · 1 August 2025
Persona vectors are activation directions for traits like sycophancy or 'evil' that let researchers monitor and steer a model's character during training. - Tracing the Thoughts of an LLM (Biology of a Large Language Model) Anthropic · 27 March 2025
Attribution-graph circuit tracing in Claude 3.5 Haiku shows shared multilingual concepts, rhyme planning ahead, parallel arithmetic paths and hallucination and jailbreak circuits. - Auditing language models for hidden objectives Anthropic · 13 March 2025
Anthropic trained a model with a hidden misaligned objective, then ran a blind auditing game with four researcher teams to test audit techniques. - Gemma Scope Google DeepMind · 31 July 2024
400+ open JumpReLU sparse autoencoders (30M+ features) for Gemma 2 2B and 9B, a 'microscope' for interpretability research. - Golden Gate Claude Anthropic · 23 May 2024
For 24 hours the public could chat with a Claude whose Golden Gate Bridge feature was turned up, a live demo that features steer behavior. - Scaling Monosemanticity Anthropic · 21 May 2024
Sparse autoencoders extract millions of interpretable features from the middle layer of Claude 3 Sonnet, including the Golden Gate Bridge feature. - Towards Monosemanticity Anthropic · 5 October 2023
Dictionary learning on a small transformer extracts 4,000+ interpretable features from a 512-neuron layer, a better unit of analysis than single neurons.
Who is working on it
- Dario Amodei, Anthropic
- Interpretability team, Anthropic
- Eric Ho, Goodfire
- Leo Gao, Dan Mossing and co-authors (sparse circuits), OpenAI
The labs with the most milestones and launches here are Anthropic (16), Google DeepMind (2), Goodfire (1), MIT Technology Review (1), Alibaba Qwen (Tongyi) (1) and OpenAI (1).
Sources
- anthropic.com/research/team/interpretability
- anthropic.com/research/natural-language-autoencoders
- anthropic.com/research/persona-vectors
- anthropic.com/research/diff-tool
- anthropic.com/research/global-workspace
- darioamodei.com/post/the-urgency-of-interpretability
- arxiv.org/abs/2511.13653
- technologyreview.com/2026/01/12/1130003/mechanistic-interpretability-ai-research-models-20
- goodfire.ai/blog/we-can-and-must-solve-alignment
This research bet was checked against its sources on 6 October 2026. How we check