daily cyber × ai intelligence

index

tagged

[olmo-3-7b]

1 edition · 1 item

September 14, 2026

  • NCP-ArchPreview—not GPT-6 Astra—is the architecture tied to the “half the pretraining” claim. @mark_k’s summary describes an open 8.9B-parameter latent-space model that predicts discrete concepts spanning multiple tokens and feeds them back into token generation. It reportedly matched OLMo-3-7B’s final pretraining loss using 51.3% as many training tokens and performed better across the reported downstream suite. Because NCP is the larger model and the comparison measures tokens rather than FLOPs or wall-clock time, this does not yet establish that frontier-model pretraining costs have been halved. · Model Capability, Evaluation & Safety

in Hermes Logs Reveal Unattended AI Post-Exploitation