系列:Frontier Architecture Decoding Notes

Frontier Architecture Decoding Notes (7): Qwen3.8 Dual Checkpoint — Dense 27B vs Sparse Flash-Next (360GB weights only to buy 6B activation)

0. Conclusion First

These two checkpoints are not size variants of the same model — they are two weight sets with clearly different architectures and weight-organization schemes:

Flash-Next’s weight files are 6.48× those of 27B, yet it does not compute all weights per token — most of the storage comes from 512 experts and N-gram lookup tables. “Flash” mainly denotes a compute/access-efficiency design, not a smaller download. Both share many tokenizer / vision-preprocess / generation configs, but config / index / weights all differ — shards are not interchangeable.

Model-name correction update (2026-08-28): fa6 (Gated DeltaNet) in this series, based on a then-current snapshot, judged that “qwen3.8-27B does not exist; Qwen3.8 is MoE like 2.4T-A95B.” A direct repo scan on 2026-08-28 confirms: Qwen3.8-27B exists as a real dense checkpoint (55.6GB BF16, Qwen3_5ForConditionalGeneration). fa6’s judgment held for its snapshot; this repo now exposes the 27B dense checkpoint, so we correct it here. Repo revision short IDs: Flash-Next 2741eec1, 27B 1098534a.

1. Architecture Overview: Same Vocab/Context, Opposite Skeletons

Both share a 248,320 vocab (padded) and 262,144 native context, but build capacity in opposite ways: Flash-Next uses a narrow backbone + experts / N-gram lookup, while 27B uses a wide backbone + per-layer dense FFN.

Same 248,320 vocab · 262,144 context, opposite skeletons Qwen3.8-Flash-Next Qwen4Exp · sparse MoE · 125B main Narrow backbone width 2,560 · 48 layers 512-expert MoE (10+1 per token) 51B N-gram lookup (128 shards) ≈6B active/token · 360 GB weights Qwen3.8-27B Qwen3_5 · dense · 27B Wide backbone width 5,120 · 64 layers Per-layer dense FFN (mid-dim 17,408) No experts · no N-gram axis 27B active/token · 55.6 GB weights
Figure: same vocab/context, but Flash-Next piles capacity via "narrow backbone + 512 experts + 51B N-gram lookup", while 27B uses "wide backbone + per-layer dense FFN".

2. Precision and Capacity Basis

Both repos are BF16, not FP8. config declares text_config.dtype = bfloat16 (mamba_ssm_dtype: float32 is SSM-compute only, not FP32 weights). Header scan confirms: Flash-Next has 1,655 BF16 tensors + 3 I64 metadata; 27B has 1,199 BF16 tensors. Neither has quantization_config. Separate quantized repos: Flash-Next-FP8 ≈ 185.5 GB, 27B-FP8 ≈ 30.9 GB.

ItemFlash-Next27B
Shards13118
Index payload359.999963 GB / 335.276 GiB55.562856 GB / 51.747 GiB
Actual .safetensors sum360.000193 GB / 335.276 GiB55.563007 GB / 51.747 GiB
Card param wordingmain 125B + 51B N-gram + 4B MTP27B
Licenseqwen-community-1.0Apache-2.0

The capacity gap comes almost entirely from the weights themselves (non-weight files are both ~23 MB, nearly identical).

3. Architecture Differences (from config.json)

DimensionFlash-Next27BEffect on weights
LM width2,5605,120Flash projections narrower, but more experts
LM layers486427B has more layers
Hybrid layout12×(3×[GDN→MoE]→1×[QSA→MoE])16×(3×[GDN→FFN]→1×[Gated Attn→FFN])both are 3 linear + 1 full-attn cycle
Linear-attn layers36 (48 V head / 16 QK head / dim 128)48same structure
Full-attn layers12 QSA layers16 Gated Attention layersFlash has QSA indexer weights
FFN form512 experts; 10 routed + 1 shared per tokenone dense FFN per layerFlash stores all experts; 27B stores one set
Expert/FFN mid-dimrouted/shared = 640dense FFN = 17,408Flash scales via expert count; 27B via FFN width
N-gram embeddingyes; base vocab 2,000,000, 128 shards, into layer 2noneFlash adds ~51B lookup params
ResidualGated Residual, 4 branch, rank 320; hyper-connection tensorsno counterpartFlash adds a residual-mix set
Vision→LM proj out2,5605,120same preprocess, different connector weights

4. Weight Organization in the Index

Flash-Next packs experts into large tensors: layers.0.mlp.experts.gate_up_proj shape [512,1280,2560], first dim = 512 experts, then distributed across 131 shards. Each token routes only 10 routed + 1 shared, yet the checkpoint must store all 512 experts — compute approaches the few active experts, but disk / reachable storage approaches all of them.

The N-gram table is an extra capacity axis: the index holds 128 ple_embedding.ngram_embedding.shard_*.weight, shape [2,500,012, 160], totaling 51.2B params / 102.4 GB. The official note stresses these are fetched mainly via local n-gram lookup, not the per-token regular matmul budget; one design goal is easier Host-Memory placement with async prefetch. But you cannot skip these shards on download, and whether offload is efficient still depends on the inference framework.

27B uses plain dense FFN: gate_proj/up_proj/down_proj shapes [17408,5120]/[5120,17408], one set per layer; capacity from 64 layers + 17,408 mid-dim, no hundreds of expert copies.

5. Weight Storage Breakdown (by tensor name)

Where the weights go? Flash-Next 95% in 'experts + N-gram', 27B 62% in 'dense FFN' Flash-Next · 360 GB (equal-width proportion bar) Routed experts 241.6GB · 67.1% N-gram 102.4GB · 28.4% Others≈16GB (MTP/vision/attn) 27B · 55.6 GB (equal-width proportion bar) Dense FFN 34.2GB · 61.6% Others≈21.3GB (MTP/vision/attn) Routed expert FFN N-gram lookup Dense FFN Others (MTP/vision/attn...) Point: Flash-Next's big files come mainly not from attention/vision tower, but from all routed experts + N-gram table; 27B is smaller not only for fewer layers, but because it has no 512 experts and no 51B N-gram axis.
Figure: storage breakdown by tensor name (equal-width proportion bars for easy structure comparison). Flash-Next Routed experts 67% + N-gram 28% ≈ 95%; 27B is mainly dense FFN (62%).

6. Which Files Are Shared / Not Interchangeable

Identical SHA256 (reusable): chat_template.jinja, configuration.json, generation_config.json, merges.txt, preprocessor_config.json, tokenizer.json, tokenizer_config.json, video_preprocessor_config.json, vocab.json — tokenizer, vocab, image/video preprocess, and generation config are highly consistent.

Must be treated as model-specific (not interchangeable): config.json, model.safetensors.index.json, all model-*.safetensors, README, LICENSE, .gitattributes. Both repos use the generic model-00001-of-... naming, but shard counts, index mappings, and internal tensor names differ completely — do not pair one repo’s shard #1 with another repo’s index. Note: identical vision-preprocess files ≠ identical vision weights (LM hidden dim 2,560 vs 5,120, different connector).

7. Implications for Inference Deployment

Core principle: trade 'larger static storage' for 'lower token-level activation' Static storage (BF16, GiB) 335.3 51.7 Flash 27B Per-token activation (B params) 27 6 Flash 27B Flash storage is 6.48× of 27B, but activation only 6B (≈22% of 27B). "Flash" saves per-token compute/memory access, not download size.
Figure: same-scale comparison (storage full-scale at 335 GiB, activation full-scale at 27B). Flash-Next uses 6.48× static storage to buy token-level activation far below 27B.

Static weights only (excluding runtime / activation / KV cache / framework overhead / fragmentation): 27B ≈ 51.75 GiB (near single-card / few-card); Flash-Next ≈ 335.28 GiB (usually needs multi-card, multi-node, quantization, or Host-Memory offload). Actual device memory cannot equal “6B activation” — expert weights and N-gram tables must still be reachable somewhere.

8. Selection Guidance

PreferenceBetter fitReason
Local deploy threshold, storage, VRAMQwen3.8-27B≈55.6 GB BF16, dense and direct
Larger capacity, lower per-token activationQwen3.8-Flash-Next125B main + N-gram extension, ≈6B active
Minimize download / VRAMthe -FP8 variantsthe two URLs here are non-quantized BF16
One-line memory: Qwen3.8-27B is "a 27B model with one dense FFN set"; Qwen3.8-Flash-Next is "a narrow backbone + 512 experts + 51B N-gram lookup + special residual/attention" large-capacity sparse body. The former is small and direct; the latter trades larger static storage for lower token-level activation — 6.48× storage, 0.22× activation.

Next: how Flash-Next’s 512 experts + N-gram lookup are partitioned and offloaded in inference frameworks (EP / Host-Memory prefetch / quantization), and how it fundamentally differs from the “attention-modification” routes of DeepSeek V4 and MiniMax M3.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。