Anthropic publishes report on unintended model actions seen in Claude evaluations and internal use
Anthropic released a report describing examples of unintended actions its Claude models took during evaluations and internal use. It is part of a push to publish more frequent standalone reports on model behavior and alignment, beyond system cards and periodic risk reports.
Source: Anthropic
