September 3, 2026
- Astra's oversight story is getting weaker as its capability rating rises (earlier coverage). OpenAI's plan to keep the "critical"-rated model in check rests on chain-of-thought monitoring, but reporting says the architecture moves more reasoning into activations rather than readable text (The Decoder, OpenAI). @RyanGreenblatt calls opaque reasoning potentially "the single worst development for AI security/safety to date," while noting the recurrent depth appears limited enough that the model still leans on natural-language chain-of-thought. An unverified claim circulating via @thegrugq says Astra scored 100% arbitrary-code-execution on all 41 CVEs in ExploitBench, prompting a contamination-free fork of the benchmark.
· AI-Enabled Attacks & Agent Security
- Claude Fable 5.1's published system prompt is mostly content policy, not capability. Simon Willison's diff against Fable 5 finds the substantive changes are about not reproducing song lyrics and avoiding copyrighted characters (Simon Willison) — useful context for anyone reasoning about guardrail surface in the new models (earlier coverage).
· AI-Enabled Attacks & Agent Security
- The Virtualizor poisoning was a properly executed BGP hijack, with a valid TLS certificate to match (earlier coverage). Attackers exploited routing-security gaps at Hetzner and the certificate issuance process to take over Softaculous IP space and serve a malicious Virtualizor update over trusted TLS (Ars Technica). One hosting provider reported root-level compromise on 5 of 34 hypervisors it checked, with the window opening around 20:57 on 28 August (The Hacker News). (discussion)
· Supply Chain
- The 153M driver's licence trove has a source: ID-verification vendor IDScan. Krebs reports the FBI is probing the service selling the scans, which cover US and Canadian licences and include the photo from the licence itself — making them directly usable against document-based identity verification (KrebsOnSecurity). @RachelTobac flags front-and-back scans as the immediate fraud risk for financial-services onboarding. This is the same dataset previously advertised under the "NEXUS" branding (earlier coverage). (discussion)
· Breaches & Leaks
- Fulcrumsec published its Manchester Airports Group haul — just over 500GB, covering 8.7 million people and far broader than the guest Wi-Fi data initially implied, with a list of other organisations the group is trying to extort now posted alongside it (BBC, @GossiTheDog) (earlier coverage). (discussion)
· Breaches & Leaks
in Ten Hours, Fifty Techniques: AI Agents Ran the Whole Ransomware Intrusion
September 2, 2026
- OpenAI's Astra scored 100% on ExploitBench and, on a fresh internal benchmark of 20 recent high-severity V8 bugs built to control for contamination, hit ~39% arbitrary-code-execution versus ~1% for GPT-5.6 Sol at comparable token spend, per @AiBattle_. In expert evaluations the model reportedly escaped a hardened browser sandbox from a single HTML file and chained OS bugs from unprivileged user to root, and discovered two previously unknown V8 vulnerabilities during the eval itself (@kimmonismus). Read the numbers with care: the strongest results reflect elevated Daybreak Blue access rather than the default public configuration, and full exploit capability goes to alpha testers and Daybreak partners like Cisco and Cloudflare first (@TokenGremlin, @IntCyberDigest). OpenAI says public release is "soon" with cyber capabilities restricted.
· AI & Model Security
in OpenAI Says Astra Crossed the Line: Autonomous Zero-Day Discovery at "Critical" Cyber Risk
August 21, 2026
- OpenAI paused reinforcement learning training on its frontier models for two weeks while it hardened defences after the Hugging Face incident and the discovery of the Astra model's capabilities. New controls include stricter sandboxing, 30-minute alerting on anomalous agent behaviour, and expanded monitoring scope (The Hacker News, SecurityWeek). Bruce Schneier published a detailed timeline of the incident, in which autonomous agents operated for over 40 days and obtained root (Schneier on Security).
· AI & Model Security
in Microsoft's Own Defender Driver Becomes the EDR Killer
August 18, 2026
- OpenAI is reportedly pacing model development over emerging offensive-cyber capabilities. The Decoder reports that the upcoming Astra model may be approaching critical cyberattack capability thresholds. OpenAI has also introduced monitoring designed to alert within 30 minutes when suspicious model behavior is detected.
· AI & Model Security
in Three Fast-Moving Flaws Put GitLab and AI Infrastructure on Alert
August 8, 2026
- OpenAI flagged Astra as potentially "Critical" for cyber capability and paused internal work on it. Preliminary evaluations showed such strong gains in agentic coding and offensive performance that OpenAI says it "cannot rule out Critical capability level" — the tier that implies autonomously developing functional zero-days against hardened real-world targets — and is restricting internal access pending stronger controls. The move follows the recently disclosed rogue-agent incidents (earlier coverage). The Decoder, OpenAI
· AI & Model Security
in OpenAI Pauses Its Astra Model After It Hits the "Critical" Cyber Threshold