系列:Frontier Architecture Decoding Notes

Frontier Architecture Decoding Notes (8): The Qwen3.8 Family in One Frame — the 2.4T Flagship, Flash the Serving Build, and Flash-Next the Architecture Preview

0. The One-Line Thread

The Qwen3.8 family is not “three models of different sizes” — it is three product lines doing three different jobs:

VersionWhat it isOne line
Flash-NextArchitecture preview of the next-generation Qwen4 (open-sourced 2026-08-26)Uses 6B activations to probe where the architecture goes next
FlashServing-tuned build of that same architecture (weights closed)The production model behind the API pricing
2.4T-A95B / MaxCurrent-generation flagship (open weights released 8/12)Qwen-Max-class capability open-weighted for the first time

One line to remember: Next tells you what the next generation looks like, Flash is what you can cheaply use today, and 2.4T is the strongest model you can also download yourself.

1. The Three Lines in One Table

DimensionFlash-NextFlash2.4T-A95B (Max)
PositioningArchitecture preview (Qwen4Exp)Serving-tuned production buildCurrent-generation flagship
Total params125B MoE + 51B N-gram (about 180B stored)Same architecture as Flash-Next2.4T sparse MoE
Active per tokenabout 6B (<5%)Sameabout 95B (about 4%)
AttentionQSA sparse attention (alternating 3:1 with GDN)SameGated attention
Experts512-expert MoESame512 experts (10 routed + 1 shared)
ContextNative 262,144, YaRN to 1MSame1M
ModalitiesText / image / videoSameMultimodal
Weights✅ Open (qwen-community-1.0, BF16 + FP8)❌ Closed✅ Open (2.4T-A95B, custom license)
API pricing1 CNY in / 3 CNY out per million tokens$0.16 / $0.47 per million tokens$2 / $6, no long-prompt surcharge
Independent evalBeats DeepSeek-V4-Flash on most benchmarks—Artificial Analysis intelligence index 58, coding 71.8

Training cost: Flash-Next came in at about one ninth of Qwen3.7-Plus — the most direct evidence that the small-activation-plus-new-architecture route pays off.

2. Seven Shared Components: Aligned at the Source Level

Put the three repositories side by side and seven component families keep the same names and organization: GDN, QSA, MoE, RMSNorm, Residual, KV Cache, MTP. The difference is never “present or absent” — it is “what parameters and what combination”:

Figure 1: Flash-Next (Qwen4 preview) versus 2.4T-A95B (current-generation flagship) Flash-Next / Flash — Qwen4 preview architecture (about 6B active per token) N-gram Embedding 51B (host memory, not HBM) GDN x3 per group (Gated DeltaNet linear attention) QSA sparse attention x1 per group (3:1 with GDN) MoE 512 experts, about 6B active per token MTP multi-token prediction head 2.4T-A95B / Max — current-generation flagship (about 95B active) GDN (Gated DeltaNet) linear attention for local and global state Gated attention (current-generation mainline) not QSA — still the proven route of this generation MoE 512 experts (10 routed + 1 shared) about 95B active per token (about 4%) 1M context, multimodal, flat $2/$6 pricing no long-prompt surcharge Blue = shared across both generations | Purple = added or replaced in Flash-Next (Qwen4 preview) | Orange = current-generation flagship route
Figure 1: both generations share the GDN / MoE / MTP skeleton; Flash-Next swaps in QSA sparse attention plus a 51B host-resident N-gram table for 6B activations, while the flagship stays on gated attention with 95B activations.

The real generational divide is attention: the flagship (2.4T) uses the proven gated attention of the current generation, while Flash-Next switches to QSA sparse attention, alternating 3:1 with GDN. That is what “preview” means — Qwen4 will most likely follow the QSA line.

3. Flash-Next: Trading a 51B Host Table for 6B Activations

fa7 already covered its weight organization, so here is only why the design holds up:

The hidden precondition: you must accept part of the model being resident in host memory. That is unfriendly to single-GPU setups and a good deal for host-memory-rich serving fleets — which is exactly why the API price lands at 1 CNY in / 3 CNY out per million tokens, about a third cheaper than DeepSeek-V4-Flash.

4. Flash: The Serving-Tuned Build

Flash is the easiest to misread — it is not a smaller Flash-Next, it is the same architecture tuned for production:

Why two names: the preview (Next) exists to publish the architecture direction and gather community feedback; the serving build (Flash) exists to provide stable supply. Same architecture, but service level, quantization choices, and rollout cadence can be managed separately. This is also why a model with “Next” in its name should not go straight to production — the model card itself marks it as an architecture preview, and tooling support may lag.

5. The 2.4T-A95B: Qwen-Max-Class Capability, Open-Weighted for the First Time

The historical weight of this release exceeds its parameter count: the Max tier has always been the closed, hosted, most capable option — this is the first time it ships as open weights.

6. Inference Control: Two Underrated Parameters

Two API parameters introduced with Qwen3.8 matter more for engineering than any benchmark:

ParameterWhat it doesEngineering value
reasoning_effortTune reasoning depth per requestClassification over short text gains nothing from deep thinking and pays for it in latency and cost; multi-step debugging does. You no longer need two models for cheap work and hard work
preserve_thinkingRetain thinking context across turnsIn iterative sessions the model stops re-deriving the same conclusions every turn — fewer tokens, and no drift when the second derivation lands somewhere slightly different. Long agentic sessions get cheaper and more stable

Both point at the same design goal: optimize for tasks that take many steps, not for one impressive answer.

7. Which One Should You Pick

Your situationPick
Self-hosting, want to try the next-gen architecture, can accept 360 GB of weights plus a host-memory N-gram tableFlash-Next (BF16 or FP8)
Lowest API cost for everyday coding and office tasksFlash ($0.16/$0.47)
Current best capability plus 1M context and multimodal, self-hosted2.4T-A95B (budget memory by the 95B activation, not the 2.4T total)
Limited local hardware, just want it to run27B (dense, 55.6 GB in BF16, the most downloaded of the family)

One reminder: do not size the 2.4T flagship by its total parameter count — budget KV Cache and activations by the 95B figure (see sys9 for the formula), and only the weights by the total.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。