Qwen3.8-Flash-Next: 125B MoE previews Qwen4 architecture, runs locally
Part of the Qwen3 8 Flash Next Release story
Qwen has released Qwen3.8-Flash-Next, an open-weights multimodal mixture-of-experts model that the company describes as "an early preview of the architecture used in Qwen4." It packs 125B total parameters but activates only 6B per token, which Simon Willison notes delivers a significant performance boost. He has been testing Unsloth quantized builds on an NVIDIA DGX Spark, including a 72.5GB UD-IQ1_S and a 78.9GB UD-Q2_K_XL, and singled out an "xhigh reasoning effort" run from the latter as his favorite so far.
Local inference results are circulating on r/LocalLLaMA. One benchmark shows the oQ4e-mtp quant hitting 45 tok/s on an M4 Max and 25 tok/s on an M2 Ultra, roughly matching Qwen3.8 27B speeds on Apple Silicon. Separately, community testers are comparing custom "Fixed" and "Sharp" prompt templates against the stock one on SWE-bench Verified using mini-SWE-agent on an RTX PRO 6000 with full 262K context and BF16 KV cache, responding to praise for the alternatives.
The model's modest active-parameter count makes a 125B-class model practical on consumer and prosumer hardware, and its role as a Qwen4 architecture preview gives local users an early look at what is coming next.
A 125B-parameter MoE with only 6B active makes near-frontier multimodal capability practical on local hardware while previewing Qwen4's architecture.
Sources
- Simon Willison2026-08-26
- r/LocalLLaMA2026-09-05
- r/LocalLLaMA2026-09-06