Alibaba Open-Sources Qwen3.8-Flash-Next: 125B MoE Model Preview of Qwen 4
Summary: Alibaba released Qwen3.8-Flash-Next as open weights — a 125B-parameter mixture-of-experts model that uses only 6B active parameters per token and serves as the developer community's first look at Qwen 4 architecture.
Key Facts
- Released August 26, 2026; licensed under Qwen Community 1.0 (commercial and research use allowed)
- 125B total parameters with only 6B active per token; a novel 51B N-gram embedding layer can sit in system RAM rather than GPU memory
- Native 262,144-token context, extendable to 1 million tokens
- Trained at one-ninth the cost of comparable models; Alibaba claims it surpasses DeepSeek-V4-Flash and Claude Opus 4.6 on coding and office benchmarks
- Framed explicitly as a Qwen 4 architecture preview, not a finished flagship
Why It Matters
Alibaba's strategy of open-sourcing an architecture preview before the full Qwen 4 launch is designed to seed the developer ecosystem and gather real-world feedback early. It continues the pattern of Chinese AI labs driving down global LLM costs with highly efficient open-weight releases — pressuring closed-model providers on pricing.
Read More
- Bloomberg — Bloomberg
- The Decoder — The Decoder