系列:每日AI热点

Daily AI Hotspot · 2026-09-16: StepAudio 3 Tops Artificial Analysis + Step-3.7-Flash Deploy Path Matures, SigLIP 2 Becomes the Default VLA Vision Tower

★ Most Worth Your Attention Today

Step’s “Step series” delivered its strongest execution combo this month: StepAudio 3’s five models (09-15) topped Artificial Analysis — 98.9% realtime rate and 1.7% ASR-WER, both global #1; Step-3.7-Flash’s open-source deploy path is now mature (vLLM stepfun37 MTP / SGLang dev EAGLE, NVFP4 4 cards). With vision+voice open-sourced in the same month and a visibly faster cadence, this is a clean execution signal for the “Step series” supply chain.

Three layers of fact:

  1. StepAudio 3 tops the charts (09-15): five models swept the Artificial Analysis speech leaderboard — 98.9% realtime rate, 1.7% ASR-WER, both global #1. Not a single metric lead; the whole lineup landed at once.
  2. Step-3.7-Flash deploy maturity: the open weights now have a clear landing path — vLLM stepfun37 (MTP) / SGLang dev-step-3.7-flash (EAGLE multi-layer draft), and it runs on NVFP4 4 cards. Edge/workstation reachability is now open, consistent with the earlier Mac Studio M4Max / DGX Spark validations.
  3. Faster execution cadence: in the same month, “vision + voice” were both open-sourced; the gap from “releasing weights” to “giving a reproducible deploy” is now just one step — the key crossing that turns a model into productive capacity.

Actionable conclusion: if you track the domestic LLM / embodied chain, Step’s moves this month are the least ambiguous execution signal — dual-line open source + dual-engine deploy readiness beats any “coming soon.” Directly relevant to my own OpenInfer / Qwen3-4B DFlash: Step-3.7-Flash has official draft paths on both vLLM and SGLang, showing that “parallel draft + domestic / low-card deploy” is already a reusable pattern.

Worth saying separately: last issue (0915) covered “auditable speculative decoding (dual-key watermark)”; this issue Step ships mature deploys on both vLLM and SGLang, effectively pushing the “domestic MoE speculative-decode stack” from demo to production. Read together, speculative decoding is now both “safe for production” (watermark) and “runnable on domestic cards” (Step deploy ready).

2. vLLM & SGLang Community Tracking

Version status: vLLM stable v0.29.0 (09-09), next candidate v0.29.1rc0 (09-13); SGLang v0.5.19 (09-05) remains the latest stable, no new tag in-window. This period’s increments come mainly from the main branch, with no formal release (same cadence as 0915).

vLLM (main · engineering maturity + multi-backend reinforcement)

New features / architecture evolution:

SGLang (v0.5.19 · default prefix cache + DFlash2 draft-net upgrade)

New features / major adaptations:

Production practice: DFlash2’s “local conv + candidate selector” is directionally clear — it aims to directly raise the acceptance length E[L] (see the deep-dive), rather than widening the draft tree. Main-branch increment, not yet in the v0.5.19 tag.

Under the Hood: DFlash2 Draft Net + Local Conv + Candidate Selector — Pushing “Acceptance Length E[L]” Up

Last issue (0915) we nailed down a formula: Speedup = E[L]×Ttarget / (Tdraft + Tverify), concluding “E[L] is the ceiling, don’t blindly widen the draft tree.” This issue DFlash2 shows a concrete way to raise that ceiling:

Qwen3.8-27B single H200 at concurrency 1 → 3.43× is the payoff of this route. It shares DNA with 0913’s “cache-boundary Pareto” and 0915’s “E[L] ceiling”: all hunt the optimal operating point on a tradeoff surface, only this issue the knob moves from “choose D” to “improve draft-net quality.”

My read: DFlash2’s direction proves it — the next phase of speculative-decode competition is not in “parallelism” but in “draft quality.” Whoever pushes E[L] higher and Tverify lower wins. For my own OpenInfer / Qwen3-4B DFlash, the next step is to empirically test whether a “local-conv-style draft net” can reproduce 3×+ on Qwen3-4B.

Standing Topics

PD disaggregation: vLLM main takes EPD three-segment split + 9 classes of KV connector (NIXL/Mooncake/LMCache) + llm-d (CNCF); SGLang side defaults the unified Radix Tree + DeepEP v2 multi-node graph capture. Crosswise: vLLM leans connector-pluggable, SGLang leans router + native unified-tree cache (#37506 unified-pool).

Architecture evolution: four speculative-decode routes (EAGLE-3 / MTP / DFlash / DSpark); this issue SGLang pushes DFlash to DFlash2 (local conv + candidate selector); sparse attention HiSparse, FP8/NVFP4 quantization (NVFP4 lm_head still a pit), multimodal diffusion keep advancing.

PyTorch vs transformers: no new “off transformers” PR this cycle. Standing position holds: pure-PyTorch in-house pays off only for flagship models; Step-3.7-Flash debugging needs transformers ≥ 5.0. Crosswise conclusion unchanged: in-house inference stacks suit high-frequency / flagship / extreme optimization, long-tail leans on transformers.

Step adaptation: StepAudio 3 (09-15, 98.9% realtime / 1.7% ASR-WER global #1) + Step-3.7-Flash deploy maturity — vLLM stepfun37 (FP8/BF16 MTP k=3; NVFP4 4 cards TP4 needs modelopt + FP8 KV cache + async-scheduling) / SGLang dev-step-3.7-flash (EAGLE multi-layer draft). Vision+voice dual-line open-sourced in the same month, cadence accelerating. Crosswise: vLLM first-class citizen > SGLang dev image.

3. AI Papers & Industry Hotspots

Today’s Focus (1 sentence)

Google DeepMind’s SigLIP 2 upgrades the vision encoder from “recognizes objects” to a “semantic + dense + multilingual” triple and has become the default vision tower of new-gen VLAs like Qwen3-VL — the very eyes that let robots go “from recognizing to doing”; while M-RoPE (Qwen3-VL) preserves image structure with 3D positional encoding and Expert Offload lets a DeepSeek V4 671B-class MoE run on limited VRAM, rounding out today’s “vision + inference landing” thread from the model layer down to the deploy layer — and on the industry side, Unitree’s IPO acceptance, Zhiyuan’s HK listing, UBTECH’s orders past ¥1.4B, XPeng’s IRON off the line, and the dual ADAS national standards are each putting a price tag on the embodied “see + act” stack.

Paper Core (SigLIP 2 · Google DeepMind, arXiv:2502.14786)

Operator Deep-Dive: M-RoPE (Qwen3-VL)

Performance Optimization: Expert Offload

Industry Hotspots (embodied companies / chain speed-dial · pinned)

⚠️ Industry developments are not investment advice.

4. The One-Line Takeaway

Step’s “Step series” delivered its strongest execution combo this month — StepAudio 3 (98.9% realtime / 1.7% ASR-WER, global #1) tops Artificial Analysis and Step-3.7-Flash is deploy-ready on both vLLM/SGLang engines (NVFP4 4 cards) — a clean execution signal for the “Step chain”; SGLang v0.5.19 makes the unified Radix Tree the default prefix cache (re-check hit rate) and DFlash2’s draft net adds local conv + candidate selector to push the acceptance length E[L] up (Qwen3.8-27B single H200 concurrency 1 at 3.43×), sharing the same formula as 0915’s “E[L] ceiling” — moving the knob from “choose D” to “improve draft-net quality”; vLLM’s main keeps Rust Core mature + dual-key watermark steady. On the papers side, SigLIP 2 (DeepMind, arXiv:2502.14786) upgrades the vision tower to a semantic+dense+multilingual triple with a four-loss recipe and is now the default eyes of new-gen VLAs like Qwen3-VL, M-RoPE preserves structure with 3D positional encoding, Expert Offload lets a DeepSeek V4 671B MoE run on limited VRAM, and Unitree’s IPO acceptance, Zhiyuan’s HK listing, UBTECH’s orders past ¥1.4B, XPeng’s IRON off the line, and the dual ADAS national standards are each putting a price tag on the embodied “see + act” stack.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。