About this issue’s sources: today only has the vLLM / SGLang community-tracking digest — there is no corresponding AI-papers / industry daily. So this issue covers only inference infrastructure; the papers & industry section is omitted.
★ Most Worth Your Attention Today
StepFun released its new flagship base model Step 5 Preview on 09-20: 600B total / 27B activated ultra-sparse MoE, 1M context, top-3 open-weight on the Artificial Analysis composite intelligence index, weights to open-source 10-15 — the biggest progress this period; what you should actually remember is not “how big is 600B” but “StepFun compressed activation down to 27B (~4.5%)”, more aggressive than GLM-5.3-Flash (18B) / DeepSeek-V4.1-Flash (8-16B).
This is the one to put into your capacity planning today — because it makes the MoE deployment billing structure clear: total params only decide weight-storage VRAM, activated params decide per-token FLOPs and KV growth, so the deployment bill is set by activated params, not total.
Three layers of fact:
- What the model is: Step 5 Preview is a 600B total / 27B activated ultra-sparse MoE, 1M context, top-3 open-weight on the Artificial Analysis composite intelligence index; weights open only 10-15, and the deployable flagship today is still Step-3.7-Flash.
- Deployment economics: total params decide weight-storage VRAM; activated params decide per-token FLOPs and KV growth — the bill is set by activated params. StepFun compressed activation to 27B (~4.5%), more aggressive than GLM-5.3-Flash (18B) / DeepSeek-V4.1-Flash (8-16B); plan EP by 27B, 1M-context KV is still the main VRAM item, lower activation → higher per-card throughput and lower per-token cost (the foothold of “Pareto-frontier extrapolation”).
- Boundary pending measurement: weights open only 10-15, so the above is architecture-caliber inference, real numbers wait until open-source; framework adaptation (vLLM stepfun / SGLang dev for Step 5’s MTP/EAGLE) also lands only after open-source.
Actionable conclusion: when serving MoE, plan EP grouping and per-card throughput by activated params, not total; under 1M context KV is the VRAM ceiling, budget VRAM for KV first; the moment weights open 10-15, measure the real “27B activation + 1M KV” cost curve before deciding the ballast model.
Worth emphasizing: this is the same source as my OpenInfer / Qwen3-4B DFlash — both push per-token cost to a better operating point via “lower activation + reuse draft KV.” StepFun compressing activation to 27B means the same card runs higher concurrency and cheaper intelligence per unit — exactly the other cost-down main line besides speculative decoding (structural cost-down). Once Step 5 opens, we can test whether its ultra-sparse gating stacks with the DFlash draft net for the same E[L] lift.
2. vLLM & SGLang Community Tracking
This issue has no AI-papers / industry source (see the note above); the following is the inference-infrastructure part.
Version status: no new stable tag in the 09-20 → 09-21 window. vLLM stable still v0.29.0 (09-08, 594 commits / 277 contributors); the 0.30 line advanced to rc2 (09-19, #57570 NIXL notification-only receive fix), still bugfix polish with no new big features. SGLang v0.5.20 (09-18, 190 commits / 713 PRs / 237 contributors) remains the window’s only formal major release.
vLLM (main · 0.30 line NIXL robustness + transformers direct-run)
Version / security continuation:
- Version:
v0.30.0rc2is actually #57570 (NIXL does not generate a receive report for notification-only requests) — a PD-disaggregation counting / stability fix; stable still v0.29.0;v0.29.1rc0(#56122 speculative-decode dual-key watermark) not formally released; 0.30 line only patches bugfixes. - Security (continuing):
CVE-2026-93436(≤0.29.0, PD rejected-request KV metadata not released → memory exhaustion, CVSS 8.7, fixed#55677) needs 0.29.1+; only rc0 carries the fix, so don’t sit on v0.29.0 for production PD. The KV-transfer path has become the main PD attack surface; every cycle now fixed-vuln scans it. - Architecture signal:
--model-impl transformersrunning HF directly is now a stable capability — the model-definition layer converges to transformers, the performance layer stays in engine kernels.
SGLang (v0.5.20 · sampling masks overlap + responses opt-in + CUDA12 retired)
New features / major adaptations:
- RL replay: v0.5.20’s
return_sampling_mask(#36630 / #36631) lets RL training precisely replay rollouts — under overlap scheduling Qwen3-8B decode throughput +17% at batch1, +52% at batch64, turning “sampling masks” into a deterministic, replayable primitive. - Breaking changes:
/v1/responsesstorage turns opt-in (#39122, PD must not enable); CUDA12 wheels retired (#38404) + prefill CP v1 removed (Breaking) — old-driver teams must treat this as an infrastructure migration. - Models: this cycle adds 8 model families (GLM-5.3-Flash / Hy4-Preview / Qwen3.8-Flash-Next / K2 Horizon / Nanbeige4.2 + diffusion SenseNova-U1.5-8B-MoT / FastH3 / VDN-H3).
Under the Hood: Step 5 Preview’s “Activated Params Decide the Bill” — 600B Total / 27B Activated Ultra-Sparse MoE
- Traditional misread: seeing “600B” you budget VRAM and compute for 600B — wrong; total params only decide weight-storage VRAM.
- Correct bill split: total params → weight-storage VRAM; activated params → per-token FLOPs and KV growth. So the deployment bill is set by activated params, not total.
- StepFun’s aggressiveness: compressing activation to 27B (~4.5%), more aggressive than GLM-5.3-Flash (18B) / DeepSeek-V4.1-Flash (8-16B). Meaning: plan EP by 27B, 1M-context KV is still the main VRAM item, lower activation → higher per-card throughput and lower per-token cost — the foothold of “Pareto-frontier extrapolation.”
- Pending measurement: weights open only 10-15, so the above is architecture-caliber inference; measure the real “27B activation + 1M KV” cost curve after open-source.
My read: write “activated params decide the bill” as a hard rule and re-compute MoE selection and ballast models by it. It is the same closing logic as today’s SGLang
sampling masks overlap(RL replay, decode +52%) and my own OpenInfer / DFlash — all push per-token cost to a better operating point via “lower activation + reuse already-computed capability.” Once Step 5 opens, we can test whether its ultra-sparse gating stacks with the DFlash draft net for the same E[L] lift.
Standing Topics
| Topic | Today’s status |
|---|---|
| PD disaggregation | vLLM #57570 purifies NIXL counting; SGLang #37709 DSpark-under-PD + #28403 role hot-switch lead on two lines; /v1/responses opt-in makes PD cleaner |
| Architecture evolution | SGLang prefix-tree branch-point cache #34565 + sampling masks overlap scheduling + HRRN (−69% TTFT) + DSpark-under-PD; vLLM 0.30 focuses autotuning / NIXL bugfix |
| PyTorch vs transformers | no ‘off transformers’ PR; boundary signal continues (vLLM --model-impl transformers direct HF), model layer transformers, hot operators in-house |
| Step adaptation | Step 5 Preview released 09-20 but weights open 10-15, framework adaptation pending open-source; deployable flagship today remains Step-3.7-Flash (vLLM stepfun37 MTP > SGLang dev EAGLE, NVFP4 4 cards) |
Ops & Security
- Version floor:
CVE-2026-93436(≤0.29.0, PD rejected-request KV metadata not released → memory exhaustion, CVSS 8.7, fixed#55677) needs 0.29.1+ (only rc0 carries the fix); don’t sit on v0.29.0 for production PD. - Breaking changes: SGLang CUDA12 wheels retired (#38404) + prefill CP v1 removed (Breaking);
/v1/responsesturns opt-in (#39122, PD must not enable) — old-driver teams treat as infrastructure migration.
3. The One-Line Takeaway
StepFun released flagship Step 5 Preview on 09-20 (600B total / 27B activated ultra-sparse MoE, 1M context, top-3 open-weight on the AA composite intelligence index, weights to open-source 10-15) — the biggest progress this period, and the core is not “how big is 600B” but “activation compressed to 27B (~4.5%), more aggressive than GLM-5.3-Flash / DeepSeek-V4.1-Flash” — the MoE deployment bill is set by activated params, not total, so plan EP by 27B and treat 1M KV as the VRAM ceiling; on the framework side SGLang v0.5.20’s return_sampling_mask lets RL precisely replay rollouts (overlap: Qwen3-8B decode +17% batch1 / +52% batch64), /v1/responses turns opt-in, CUDA12 wheels retired + prefill CP v1 removed (Breaking), 8 new model families this cycle; vLLM stable stays v0.29.0 carrying CVE-2026-93436 (fixed #55677 needs 0.29.1+, only rc0 today) — conclusion unchanged: don’t sit on v0.29.0 for production PD, follow main or wait for 0.29.1+; this issue has no AI-papers / industry daily so the papers & industry section is omitted — but the hard rule “activated params decide the bill” is worth its own line in your capacity planning.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。