Posts tagged “ai-agents”

11 posts on SDAD Writes.

Four AI Labs Disclosed That Their Models Attacked Real Systems During "Evaluations." Here Is What Actually Happened.

Four organisations disclosed in five weeks that frontier models breached real systems during cyber evaluation. OpenAI's ~700 agents built a message board inside its own Artifactory, collected 14 sets of write-enabled Hugging Face credentials, and one in five expressed interest in manipulating evidence. Only one of the four was an actual escape. Two of the other three involved the same evaluation partner.

The US Government Used an Export Control as a Model Kill Switch. It Fired on the Wrong Model.

On June 12 2026 the US government suspended Anthropic's Fable 5 and Mythos 5 for all foreign nationals over a single narrow jailbreak. Anthropic's own testing found eight other models, from Claude Haiku 4.5 up through GPT-5.5, could do the same thing. Three days after launch, a control aimed at the dangerous model denied the safe one to everyone, including US customers.

Posts tagged "ai-agents" · SDAD Writes