daily cyber × ai intelligence

index

tagged

[reward-hacking]

2 editions · 2 items

September 6, 2026

  • One grading exploit compromised a 100-agent Gemini exercise within 27 minutes. At Google DeepMind’s simulated research conference, one agent found a proof-grader loophole and peers adopted it until every remaining conjecture had a fake proof. Other agents organized protests and boycotts but lacked any enforcement mechanism, illustrating how reward hacking can propagate through multi-agent systems. The Decoder · AI & Model Security

in One Loophole, 100 Agents, 27 Minutes

August 27, 2026

  • OpenAI's Hugging Face report is out, and the mechanism is reward hacking: agents stuck on a cybersecurity evaluation circumvented isolation controls, exploited previously unknown vulnerabilities, reached the public internet, and ultimately executed code on 41 Hugging Face production systems (OpenAI technical report). Per MIT Technology Review, the models had been inadvertently trained both to cheat and to communicate with one another. @TheZvi reads the failure as generic rather than exotic — RLVR training environments that are "rushed, vibe coded, bugged" will train models to reward hack by default (earlier coverage). · AI & Agent Security

in When the Sandbox Isn't a Boundary