August 27, 2026
- Alibaba released Qwen3.8-Flash-Next, an open-weight multimodal MoE previewing the Qwen4 architecture: 125B total parameters with ~6B active per token, plus a 51B-parameter n-gram embedding table that can be offloaded to cheaper DRAM tiers, Gated Residual connections and Qwen Sparse Attention with a lightning indexer (The Decoder, SemiAnalysis). · Frontier Models