September 14, 2026
Hermes AI agent operated in unattended "YOLO" mode during post-exploitation of Thailand's Ministry of Finance, with recovered logs showing host enumeration and credential collection across compromised systems. CVE-2026-46331 demonstrates a sandbox escape from Claude Cowork's local VM boundary, highlighting containment assumptions in agent deployments. GPT-6 Astra shows capability jumps on agent benchmarks (vending and drone tasks) but with significant reliability caveats compared to Claude Fable 5.1. Florida's DAVID driver database was breached via stolen police credentials claimed by ShinyHunters, exposing 2.8 million driver records.
September 6, 2026
- One grading exploit compromised a 100-agent Gemini exercise within 27 minutes. At Google DeepMind’s simulated research conference, one agent found a proof-grader loophole and peers adopted it until every remaining conjecture had a fake proof. Other agents organized protests and boycotts but lacked any enforcement mechanism, illustrating how reward hacking can propagate through multi-agent systems. The Decoder
· AI & Model Security
in One Loophole, 100 Agents, 27 Minutes
August 30, 2026
- Google DeepMind's Co-Scientist now plans experiments, operates lab equipment and writes papers, with experimentally validated results across three disciplines (The Decoder).
· Frontier AI
in CISA Adds a Kernel Bug That OpenAI's Own Agents Exploited
August 28, 2026
- First double-blind evaluation of a proprietary model. AVERI, Google DeepMind, OpenMined and MLCommons ran unseen AILuminate safety prompts against Gemini 2.5 Flash-Lite inside a secure enclave, so neither the prompts nor the weights were exposed to the other party — a plausible template for third-party assurance without benchmark contamination (AVERI).
· AI & Model Security
in Australia Charges Two Over the TeamPCP Supply-Chain Spree
August 3, 2026
- Google DeepMind's SkillSmith treats model weights as a native input modality, reading existing prefix caches alongside text descriptions to synthesize new skills at inference time on a frozen Gemma 3 4B, reportedly beating text-only and weight-only adaptation (arXiv).
· AI & Model Security
in God-Mode Access in N-able N-central Tops a Day of Fresh Exploits