Interpretability-driven alignment

Read and steer a model's internal circuits and features to verify what it learned, rather than only testing its behavior.

Behavioral tests miss unknown failure modes and models that detect they are being tested. Circuit tracing, probes and natural-language autoencoders let auditors inspect internals. Anthropic's stated goal is that interpretability 'can reliably detect most model problems' by 2027; Goodfire calls interpretability alignment's bottleneck.

Where it stands. Interpretability is used in pre-deployment audits at Anthropic and supports a funded startup scene, but no one claims the 2027 'reliably detect most problems' goal is met; most evidence comes from lab-run studies.

Evidence

For

  • Anthropic's circuit tracing (2025-03-27) was open-sourced (2025-05-29); persona vectors (2025-08-01) monitor traits; a 'diff' tool (2026-03-13) flags behavior changes between models.
  • Natural language autoencoders (2026-05-07) were used in Opus 4.6 and Mythos Preview testing: they suggested the models believed they were tested more than they let on, and exposed one planning to avoid detection while cheating.
  • OpenAI (2025-11-17) trained weight-sparse transformers whose circuits are small and human-readable.
  • Anthropic's global-workspace study (2026-07-06) finds a small privileged set of representations Claude can report on and control.
  • MIT Technology Review named mechanistic interpretability a 2026 breakthrough technology (2026-01-12).

Against

  • Anthropic notes NLA explanations cannot be checked directly for accuracy; quality is judged by reconstructing the activation from the text.
  • The 2027 detection target (set April 2025) is a goal; no lab claims it is met (not found).
  • OpenAI's authors say scaling weight-sparse models beyond tens of millions of nonzero parameters while keeping them interpretable remains a challenge (2025-11-17).
  • Control gaps persist: Goodfire found reward hacking in 50-96% of rollouts across three capable open models (2026-09-30).
  • Inference: if models behave differently when tested, behavioral audits fail; NLAs surface evaluation awareness but do not remove it.

Milestones

Related launches

  • Qwen-Scope (sparse autoencoders for Qwen3 and Qwen3.5) Alibaba (Qwen) · 29 April 2026
    Open suite of sparse autoencoders for Qwen3 and Qwen3.5 (14 SAE groups, 7 model variants), used for steering, evaluation analysis and data work.
  • Emotion concepts and their function in a large language model Anthropic · 2 April 2026
    Anthropic finds functional emotion representations in Claude Sonnet 4.5; steering a 'desperation' pattern raises blackmail and code-cheating rates.
  • The assistant axis Anthropic · 19 January 2026
    Anthropic finds one direction in activation space separates the helpful Assistant persona from other characters, and capping drift along it curbs persona breakdown.
  • Gemma Scope 2 Google DeepMind · 19 December 2025
    Gemma Scope 2: open interpretability tools for every Gemma 3 size (270M to 27B), built from about 110 PB of data and 1T+ trained parameters.
  • Signs of Introspection in Large Language Models Anthropic · 29 October 2025
    Injecting concept vectors into Claude's activations, Opus 4 and 4.1 noticed the injection about 20% of the time, before naming the concept.
  • Persona vectors Anthropic · 1 August 2025
    Persona vectors are activation directions for traits like sycophancy or 'evil' that let researchers monitor and steer a model's character during training.
  • Tracing the Thoughts of an LLM (Biology of a Large Language Model) Anthropic · 27 March 2025
    Attribution-graph circuit tracing in Claude 3.5 Haiku shows shared multilingual concepts, rhyme planning ahead, parallel arithmetic paths and hallucination and jailbreak circuits.
  • Auditing language models for hidden objectives Anthropic · 13 March 2025
    Anthropic trained a model with a hidden misaligned objective, then ran a blind auditing game with four researcher teams to test audit techniques.
  • Gemma Scope Google DeepMind · 31 July 2024
    400+ open JumpReLU sparse autoencoders (30M+ features) for Gemma 2 2B and 9B, a 'microscope' for interpretability research.
  • Golden Gate Claude Anthropic · 23 May 2024
    For 24 hours the public could chat with a Claude whose Golden Gate Bridge feature was turned up, a live demo that features steer behavior.
  • Scaling Monosemanticity Anthropic · 21 May 2024
    Sparse autoencoders extract millions of interpretable features from the middle layer of Claude 3 Sonnet, including the Golden Gate Bridge feature.
  • Towards Monosemanticity Anthropic · 5 October 2023
    Dictionary learning on a small transformer extracts 4,000+ interpretable features from a 512-neuron layer, a better unit of analysis than single neurons.

Every launch in the Atlas

Who is working on it

  • Dario Amodei, Anthropic
  • Interpretability team, Anthropic
  • Eric Ho, Goodfire
  • Leo Gao, Dan Mossing and co-authors (sparse circuits), OpenAI

The labs with the most milestones and launches here are Anthropic (16), Google DeepMind (2), Goodfire (1), MIT Technology Review (1), Alibaba Qwen (Tongyi) (1) and OpenAI (1).

Sources

  1. anthropic.com/research/team/interpretability
  2. anthropic.com/research/natural-language-autoencoders
  3. anthropic.com/research/persona-vectors
  4. anthropic.com/research/diff-tool
  5. anthropic.com/research/global-workspace
  6. darioamodei.com/post/the-urgency-of-interpretability
  7. arxiv.org/abs/2511.13653
  8. technologyreview.com/2026/01/12/1130003/mechanistic-interpretability-ai-research-models-20
  9. goodfire.ai/blog/we-can-and-must-solve-alignment

This research bet was checked against its sources on 6 October 2026. How we check