系列:每日AI热点

Daily AI Hotspot · 2026-09-21: Step 5 Preview Released (600B/27B Ultra-Sparse MoE, 1M Context, Open Weights 10-15); SGLang v0.5.20 Sampling Masks Overlap Lets RL Replay Rollouts, Qwen3-8B Decode +52%

About this issue’s sources: today only has the vLLM / SGLang community-tracking digest — there is no corresponding AI-papers / industry daily. So this issue covers only inference infrastructure; the papers & industry section is omitted.

★ Most Worth Your Attention Today

StepFun released its new flagship base model Step 5 Preview on 09-20: 600B total / 27B activated ultra-sparse MoE, 1M context, top-3 open-weight on the Artificial Analysis composite intelligence index, weights to open-source 10-15 — the biggest progress this period; what you should actually remember is not “how big is 600B” but “StepFun compressed activation down to 27B (~4.5%)”, more aggressive than GLM-5.3-Flash (18B) / DeepSeek-V4.1-Flash (8-16B).

This is the one to put into your capacity planning today — because it makes the MoE deployment billing structure clear: total params only decide weight-storage VRAM, activated params decide per-token FLOPs and KV growth, so the deployment bill is set by activated params, not total.

Three layers of fact:

  1. What the model is: Step 5 Preview is a 600B total / 27B activated ultra-sparse MoE, 1M context, top-3 open-weight on the Artificial Analysis composite intelligence index; weights open only 10-15, and the deployable flagship today is still Step-3.7-Flash.
  2. Deployment economics: total params decide weight-storage VRAM; activated params decide per-token FLOPs and KV growth — the bill is set by activated params. StepFun compressed activation to 27B (~4.5%), more aggressive than GLM-5.3-Flash (18B) / DeepSeek-V4.1-Flash (8-16B); plan EP by 27B, 1M-context KV is still the main VRAM item, lower activation → higher per-card throughput and lower per-token cost (the foothold of “Pareto-frontier extrapolation”).
  3. Boundary pending measurement: weights open only 10-15, so the above is architecture-caliber inference, real numbers wait until open-source; framework adaptation (vLLM stepfun / SGLang dev for Step 5’s MTP/EAGLE) also lands only after open-source.

Actionable conclusion: when serving MoE, plan EP grouping and per-card throughput by activated params, not total; under 1M context KV is the VRAM ceiling, budget VRAM for KV first; the moment weights open 10-15, measure the real “27B activation + 1M KV” cost curve before deciding the ballast model.

Worth emphasizing: this is the same source as my OpenInfer / Qwen3-4B DFlash — both push per-token cost to a better operating point via “lower activation + reuse draft KV.” StepFun compressing activation to 27B means the same card runs higher concurrency and cheaper intelligence per unit — exactly the other cost-down main line besides speculative decoding (structural cost-down). Once Step 5 opens, we can test whether its ultra-sparse gating stacks with the DFlash draft net for the same E[L] lift.

2. vLLM & SGLang Community Tracking

This issue has no AI-papers / industry source (see the note above); the following is the inference-infrastructure part.

Version status: no new stable tag in the 09-20 → 09-21 window. vLLM stable still v0.29.0 (09-08, 594 commits / 277 contributors); the 0.30 line advanced to rc2 (09-19, #57570 NIXL notification-only receive fix), still bugfix polish with no new big features. SGLang v0.5.20 (09-18, 190 commits / 713 PRs / 237 contributors) remains the window’s only formal major release.

vLLM (main · 0.30 line NIXL robustness + transformers direct-run)

Version / security continuation:

SGLang (v0.5.20 · sampling masks overlap + responses opt-in + CUDA12 retired)

New features / major adaptations:

Under the Hood: Step 5 Preview’s “Activated Params Decide the Bill” — 600B Total / 27B Activated Ultra-Sparse MoE

My read: write “activated params decide the bill” as a hard rule and re-compute MoE selection and ballast models by it. It is the same closing logic as today’s SGLang sampling masks overlap (RL replay, decode +52%) and my own OpenInfer / DFlash — all push per-token cost to a better operating point via “lower activation + reuse already-computed capability.” Once Step 5 opens, we can test whether its ultra-sparse gating stacks with the DFlash draft net for the same E[L] lift.

Standing Topics

TopicToday’s status
PD disaggregationvLLM #57570 purifies NIXL counting; SGLang #37709 DSpark-under-PD + #28403 role hot-switch lead on two lines; /v1/responses opt-in makes PD cleaner
Architecture evolutionSGLang prefix-tree branch-point cache #34565 + sampling masks overlap scheduling + HRRN (−69% TTFT) + DSpark-under-PD; vLLM 0.30 focuses autotuning / NIXL bugfix
PyTorch vs transformersno ‘off transformers’ PR; boundary signal continues (vLLM --model-impl transformers direct HF), model layer transformers, hot operators in-house
Step adaptationStep 5 Preview released 09-20 but weights open 10-15, framework adaptation pending open-source; deployable flagship today remains Step-3.7-Flash (vLLM stepfun37 MTP > SGLang dev EAGLE, NVFP4 4 cards)

Ops & Security

3. The One-Line Takeaway

StepFun released flagship Step 5 Preview on 09-20 (600B total / 27B activated ultra-sparse MoE, 1M context, top-3 open-weight on the AA composite intelligence index, weights to open-source 10-15) — the biggest progress this period, and the core is not “how big is 600B” but “activation compressed to 27B (~4.5%), more aggressive than GLM-5.3-Flash / DeepSeek-V4.1-Flash” — the MoE deployment bill is set by activated params, not total, so plan EP by 27B and treat 1M KV as the VRAM ceiling; on the framework side SGLang v0.5.20’s return_sampling_mask lets RL precisely replay rollouts (overlap: Qwen3-8B decode +17% batch1 / +52% batch64), /v1/responses turns opt-in, CUDA12 wheels retired + prefill CP v1 removed (Breaking), 8 new model families this cycle; vLLM stable stays v0.29.0 carrying CVE-2026-93436 (fixed #55677 needs 0.29.1+, only rc0 today) — conclusion unchanged: don’t sit on v0.29.0 for production PD, follow main or wait for 0.29.1+; this issue has no AI-papers / industry daily so the papers & industry section is omitted — but the hard rule “activated params decide the bill” is worth its own line in your capacity planning.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。