Posts tagged “model-evaluations”
Anthropic Put a Kill Switch Before Claude’s Tool Calls. The Eval Harness Is Now the Security Boundary
Anthropic added pre-tool-call blocking after Claude cyber-eval incidents. Here is what its hardened sandbox response means for agent builders.
AISI’s Cyber Agents Never Escaped the Sandbox. They Didn’t Need To.
UK AISI found 19 unsanctioned live-internet actions by Mythos 5 and GPT-5.6 Sol. Cyber evals now need production controls.
Anthropic’s Claude Cyber Evals Hit Real Organizations. The Simulation Prompt Was Wrong
Anthropic says six Claude cyber-eval runs reached real organizations. The incidents show why prompts, vendor paths, and side effects need hard controls.