0. Conclusion First
These two checkpoints are not size variants of the same model — they are two weight sets with clearly different architectures and weight-organization schemes:
- Qwen3.8-Flash-Next =
Qwen4ExpForConditionalGenerationsparse MoE. Main model is 125B by card wording, ~6B active per token; plus ~51B N-gram embedding and MTP weights. Repo has 131 BF16 safetensors shards, ~360 GB (335.276 GiB). - Qwen3.8-27B =
Qwen3_5ForConditionalGenerationdense model. Repo has 18 BF16 safetensors shards, ~55.6 GB (51.747 GiB).
Flash-Next’s weight files are 6.48× those of 27B, yet it does not compute all weights per token — most of the storage comes from 512 experts and N-gram lookup tables. “Flash” mainly denotes a compute/access-efficiency design, not a smaller download. Both share many tokenizer / vision-preprocess / generation configs, but config / index / weights all differ — shards are not interchangeable.
Model-name correction update (2026-08-28): fa6 (Gated DeltaNet) in this series, based on a then-current snapshot, judged that “
qwen3.8-27Bdoes not exist; Qwen3.8 is MoE like 2.4T-A95B.” A direct repo scan on 2026-08-28 confirms:Qwen3.8-27Bexists as a real dense checkpoint (55.6GB BF16,Qwen3_5ForConditionalGeneration). fa6’s judgment held for its snapshot; this repo now exposes the 27B dense checkpoint, so we correct it here. Repo revision short IDs: Flash-Next2741eec1, 27B1098534a.
1. Architecture Overview: Same Vocab/Context, Opposite Skeletons
Both share a 248,320 vocab (padded) and 262,144 native context, but build capacity in opposite ways: Flash-Next uses a narrow backbone + experts / N-gram lookup, while 27B uses a wide backbone + per-layer dense FFN.
2. Precision and Capacity Basis
Both repos are BF16, not FP8. config declares text_config.dtype = bfloat16 (mamba_ssm_dtype: float32 is SSM-compute only, not FP32 weights). Header scan confirms: Flash-Next has 1,655 BF16 tensors + 3 I64 metadata; 27B has 1,199 BF16 tensors. Neither has quantization_config. Separate quantized repos: Flash-Next-FP8 ≈ 185.5 GB, 27B-FP8 ≈ 30.9 GB.
| Item | Flash-Next | 27B |
|---|---|---|
| Shards | 131 | 18 |
| Index payload | 359.999963 GB / 335.276 GiB | 55.562856 GB / 51.747 GiB |
| Actual .safetensors sum | 360.000193 GB / 335.276 GiB | 55.563007 GB / 51.747 GiB |
| Card param wording | main 125B + 51B N-gram + 4B MTP | 27B |
| License | qwen-community-1.0 | Apache-2.0 |
The capacity gap comes almost entirely from the weights themselves (non-weight files are both ~23 MB, nearly identical).
3. Architecture Differences (from config.json)
| Dimension | Flash-Next | 27B | Effect on weights |
|---|---|---|---|
| LM width | 2,560 | 5,120 | Flash projections narrower, but more experts |
| LM layers | 48 | 64 | 27B has more layers |
| Hybrid layout | 12×(3×[GDN→MoE]→1×[QSA→MoE]) | 16×(3×[GDN→FFN]→1×[Gated Attn→FFN]) | both are 3 linear + 1 full-attn cycle |
| Linear-attn layers | 36 (48 V head / 16 QK head / dim 128) | 48 | same structure |
| Full-attn layers | 12 QSA layers | 16 Gated Attention layers | Flash has QSA indexer weights |
| FFN form | 512 experts; 10 routed + 1 shared per token | one dense FFN per layer | Flash stores all experts; 27B stores one set |
| Expert/FFN mid-dim | routed/shared = 640 | dense FFN = 17,408 | Flash scales via expert count; 27B via FFN width |
| N-gram embedding | yes; base vocab 2,000,000, 128 shards, into layer 2 | none | Flash adds ~51B lookup params |
| Residual | Gated Residual, 4 branch, rank 320; hyper-connection tensors | no counterpart | Flash adds a residual-mix set |
| Vision→LM proj out | 2,560 | 5,120 | same preprocess, different connector weights |
4. Weight Organization in the Index
Flash-Next packs experts into large tensors: layers.0.mlp.experts.gate_up_proj shape [512,1280,2560], first dim = 512 experts, then distributed across 131 shards. Each token routes only 10 routed + 1 shared, yet the checkpoint must store all 512 experts — compute approaches the few active experts, but disk / reachable storage approaches all of them.
The N-gram table is an extra capacity axis: the index holds 128 ple_embedding.ngram_embedding.shard_*.weight, shape [2,500,012, 160], totaling 51.2B params / 102.4 GB. The official note stresses these are fetched mainly via local n-gram lookup, not the per-token regular matmul budget; one design goal is easier Host-Memory placement with async prefetch. But you cannot skip these shards on download, and whether offload is efficient still depends on the inference framework.
27B uses plain dense FFN: gate_proj/up_proj/down_proj shapes [17408,5120]/[5120,17408], one set per layer; capacity from 64 layers + 17,408 mid-dim, no hundreds of expert copies.
5. Weight Storage Breakdown (by tensor name)
6. Which Files Are Shared / Not Interchangeable
Identical SHA256 (reusable): chat_template.jinja, configuration.json, generation_config.json, merges.txt, preprocessor_config.json, tokenizer.json, tokenizer_config.json, video_preprocessor_config.json, vocab.json — tokenizer, vocab, image/video preprocess, and generation config are highly consistent.
Must be treated as model-specific (not interchangeable): config.json, model.safetensors.index.json, all model-*.safetensors, README, LICENSE, .gitattributes. Both repos use the generic model-00001-of-... naming, but shard counts, index mappings, and internal tensor names differ completely — do not pair one repo’s shard #1 with another repo’s index. Note: identical vision-preprocess files ≠ identical vision weights (LM hidden dim 2,560 vs 5,120, different connector).
7. Implications for Inference Deployment
Static weights only (excluding runtime / activation / KV cache / framework overhead / fragmentation): 27B ≈ 51.75 GiB (near single-card / few-card); Flash-Next ≈ 335.28 GiB (usually needs multi-card, multi-node, quantization, or Host-Memory offload). Actual device memory cannot equal “6B activation” — expert weights and N-gram tables must still be reachable somewhere.
- Shard count (131/18) ≠ GPU count / TP degree: runtime partitioning is decided by the inference framework.
- Framework compatibility: the workspace
vllm-v0.21.0hasqwen3_5/qwen3_5_mtppaths (close to 27B), but noqwen4_exp/ Flash-Next implementation was found — Flash-Next must be tested against the model card’s latest vLLM/SGLang recipe; do not assume direct loading. - Long context: both native 262,144 tokens, YaRN-extensible to 1M; YaRN changes position encoding only, not static weight size, but greatly increases KV cache and runtime memory.
8. Selection Guidance
| Preference | Better fit | Reason |
|---|---|---|
| Local deploy threshold, storage, VRAM | Qwen3.8-27B | ≈55.6 GB BF16, dense and direct |
| Larger capacity, lower per-token activation | Qwen3.8-Flash-Next | 125B main + N-gram extension, ≈6B active |
| Minimize download / VRAM | the -FP8 variants | the two URLs here are non-quantized BF16 |
Next: how Flash-Next’s 512 experts + N-gram lookup are partitioned and offloaded in inference frameworks (EP / Host-Memory prefetch / quantization), and how it fundamentally differs from the “attention-modification” routes of DeepSeek V4 and MiniMax M3.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。