September 14, 2026
- GPT-6 Astra’s new agent benchmarks show a capability jump with large reliability caveats. Across six Vending-Bench runs, each simulating a year from a $500 starting balance, GPT-6 Astra averaged a final balance of $15,515 versus $5,422 for Claude Fable 5.1, according to The Decoder’s report on Andon Labs’ tests. On Drone-Bench, Astra’s best attempts beat the human-AI baseline on all five subtasks, including code for finding and following a specified person, but overall success remained unreliable. A separate robotics test had Astra complete 7 of 100 dual-arm tasks; MolmoAct2 completed 0 of the same 100. These were controlled evaluations, not live business or surveillance deployments, and they materially extend the initial Astra coverage (earlier coverage). · Model Capability, Evaluation & Safety