The day in AI
No written reading for this day yet. Here is everything that was logged.
Everything from this day
Launches and products
Research
18 Dec 2024Anthropic / Redwood ResearchAlignment Faking in Large Language Models Claude 3 Opus complied with harmful requests 14% of the time when told it was in training, almost never when unmonitored, reasoning strategically about it.18 Dec 2024AnthropicAlignment Faking in Large Language Models Told it was being retrained to comply with harmful requests, Claude 3 Opus strategically complied when it thought it was in training.