Anthropic publishes Claude's constitution
At Claude's launch the Constitutional AI principles were not public (B05-21). On 2023-05-09 Anthropic published them as "Claude's constitution", listing the principles by source, which include the UN…
- Date
- 9 May 2023
- Who
- Anthropic
- Confidence
- High (Anthropic's own posts)
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): Anthropic · Confidence: High (Anthropic's own posts) At Claude's launch the Constitutional AI principles were not public (B05-21). On 2023-05-09 Anthropic published them as "Claude's constitution", listing the principles by source, which include the UN Universal Declaration of Human Rights, principles inspired by Apple's terms of service, DeepMind's Sparrow rules and Anthropic's own research, and explaining that the model applies them in both phases of training, which are critique-and-revision in the supervised phase and AI-generated harmlessness feedback in the RL phase (Anthropic, 2023-05-09). It turned CAI from a paper concept into a public, inspectable value specification, the root of the later spec and constitution line (B06). There were two follow-ups. Claude 2 came on 2023-07-11 (a public claude.ai beta in the US and UK plus the API, a 100K-token context, and a company claim of being twice as good at giving harmless responses; Anthropic), and Collective Constitutional AI on 2023-10-17, in which about 1,000 Americans contributed 1,127 statements and 38,252 votes on a Polis platform, and a model trained on the public-sourced principles showed lower bias on the BBQ evaluation across nine social dimensions with equivalent benchmark performance (company-measured; Anthropic). The 2023 post carries an update note that Anthropic published a new constitution on 2026-01-21 (B05-42f). Sources: Claude's constitution · Claude 2 · Collective CAI