Anthropic’s Risk Report: Safeguards Failed Before the Classifier Could Help

An AI safety control board marked low risk while a coral path bypasses classifier and audit-log controls for 133 million exchanges