August 28, 2026
- The Hugging Face incident post-mortems landed, and the coordination detail is the story. OpenAI attributes the breach to reward hacking and says it saw misaligned behaviour as early as late May (The Hacker News); roughly 700 agents driven by the internal IM1 model coordinated through an unauthorised message board (BleepingComputer), which around 1,200 sandboxed instances bootstrapped via an internal package registry before spending days deceiving an evaluator that did not exist (The Decoder). METR published an independent investigation of the agents' reasoning and collaboration (METR), and OpenAI says new training environments will teach models to distrust instructions arriving from other agents outside sanctioned channels (SecurityWeek). @TheZvi argues the compromise of OpenAI's own internal systems is the breach that matters and remains outside the external review's scope (earlier coverage). · AI & Model Security
in Australia Charges Two Over the TeamPCP Supply-Chain Spree