September 16, 2026
- RSIAgent accumulates environment knowledge without changing its base models’ weights. Using Kimi-K3 and GLM-5.3, the framework selects experiments, verifies outcomes and stores action-condition-outcome relationships for later tasks; its claimed benchmark wins over GPT-6 Astra remain author-reported (@huang_biwei). @RitwikSrivast11 calls weight-frozen RSI “a harness that keeps score,” distinguishing test-time memory from model self-modification.
· AI & Agent Security
in CVE-2026-76461 Gives Remote Attackers Root on Cisco Email Gateways
September 14, 2026
- GPT-6 Astra’s new agent benchmarks show a capability jump with large reliability caveats. Across six Vending-Bench runs, each simulating a year from a $500 starting balance, GPT-6 Astra averaged a final balance of $15,515 versus $5,422 for Claude Fable 5.1, according to The Decoder’s report on Andon Labs’ tests. On Drone-Bench, Astra’s best attempts beat the human-AI baseline on all five subtasks, including code for finding and following a specified person, but overall success remained unreliable. A separate robotics test had Astra complete 7 of 100 dual-arm tasks; MolmoAct2 completed 0 of the same 100. These were controlled evaluations, not live business or surveillance deployments, and they materially extend the initial Astra coverage (earlier coverage).
· Model Capability, Evaluation & Safety
- NCP-ArchPreview—not GPT-6 Astra—is the architecture tied to the “half the pretraining” claim. @mark_k’s summary describes an open 8.9B-parameter latent-space model that predicts discrete concepts spanning multiple tokens and feeds them back into token generation. It reportedly matched OLMo-3-7B’s final pretraining loss using 51.3% as many training tokens and performed better across the reported downstream suite. Because NCP is the larger model and the comparison measures tokens rather than FLOPs or wall-clock time, this does not yet establish that frontier-model pretraining costs have been halved.
· Model Capability, Evaluation & Safety
in Hermes Logs Reveal Unattended AI Post-Exploitation
September 13, 2026
- Datasette shipped 1.0a39 and 0.65.4 security releases after an audit run with Claude Fable 5.1, GPT-5.6 Sol and GPT-6 Astra turned up a range of bugs; public instances should upgrade (Datasette).
· AI & Model Security
in Artifactory Chains Give Attackers Admin in Under Five Minutes
September 9, 2026
- The Astra oversight debate turned into cross-lab benchmarking. BleepingComputer reports OpenAI's position that GPT-6 Astra can autonomously find zero-days but is harder to monitor (BleepingComputer); Anthropic's Boris Cherny publicly scored the new model as "roughly on par with Gemini Flash and Opus 4.8 on prompt injection risk" and claimed Anthropic "solved prompt injection in practice for Claude models about two months ago" — a claim no third party has verified (earlier coverage).
· AI & Model Security
in One Phone Call, Zero Clicks: A WeChat Worm Crossed iOS and Android
September 7, 2026
- GPT-6 Astra is reaching $20 Plus subscribers, showing up in ChatGPT Work before the regular chat model picker; it is already generally available to Pro, Enterprise and Business Premium users in Work and Codex and in the API, with no date for free users (BleepingComputer). On the monitorability question, @RyanGreenblatt walked back his own reading that
reasoning=None was pulled for safety reasons — "I now think I was reading into this too much" — after being told it was removed simply because low dominated it on usefulness.
· AI & Model Security
in The Diff Is the Disclosure: MikroTik's Silent Patch Comes Apart
September 6, 2026
- GPT-6 Astra’s reported 99.99% direct-injection block rate does not carry over to indirect attacks. Instructions hidden inside documents succeeded in 8.5% of scenarios, compared with 4.8% for Claude Opus 5. That is the more relevant exposure for autonomous agents ingesting untrusted files and web content. The Decoder adds security detail to the model release (earlier coverage)
· AI & Model Security
in One Loophole, 100 Agents, 27 Minutes
September 5, 2026
- GPT-6 Astra is out, scoring 100% on ExploitBench with OpenAI blocking PoC exploit generation requests at the product layer — the shipping counterpart to the "Critical" cybersecurity rating under its Preparedness Framework (The Hacker News, earlier coverage). Benchmarks disagree sharply — Epoch AI puts it clearly in front, Artificial Analysis rates it no better than its predecessor — but its ARC-AGI-3 efficiency beat the average human for the first time, pulling Chollet's AGI forecast forward. @TheZvi flags the practical catch for anyone relying on oversight: the chain-of-thought is now harder to monitor and easier for the model to hide things in.
· Agentic AI & Model Security
in 18,000 Posts on a Dead German Wiki: OpenAI's Agents Were Trading Sandbox Escapes in May
September 4, 2026
- OpenAI is framing GPT-6 Astra as the start of the "AGI era", with president Greg Brockman making the call; the model tops math, coding and cybersecurity benchmarks, is the first rated "critical" under OpenAI's preparedness framework, and found two previously unknown zero-days during testing (The Decoder, SecurityWeek) (earlier coverage). @TheZvi flags the uncomfortable read on its weak monitorability: if that isn't a mistake but simply how smarter models behave, it's the worse outcome.
· Frontier AI
in Malware That Gaslights the AI Analyst