★ Most Worth Your Attention Today
Step’s “Step series” delivered its strongest execution combo this month: StepAudio 3’s five models (09-15) topped Artificial Analysis — 98.9% realtime rate and 1.7% ASR-WER, both global #1; Step-3.7-Flash’s open-source deploy path is now mature (vLLM stepfun37 MTP / SGLang dev EAGLE, NVFP4 4 cards). With vision+voice open-sourced in the same month and a visibly faster cadence, this is a clean execution signal for the “Step series” supply chain.
Three layers of fact:
- StepAudio 3 tops the charts (09-15): five models swept the Artificial Analysis speech leaderboard — 98.9% realtime rate, 1.7% ASR-WER, both global #1. Not a single metric lead; the whole lineup landed at once.
- Step-3.7-Flash deploy maturity: the open weights now have a clear landing path — vLLM
stepfun37(MTP) / SGLangdev-step-3.7-flash(EAGLE multi-layer draft), and it runs on NVFP4 4 cards. Edge/workstation reachability is now open, consistent with the earlier Mac Studio M4Max / DGX Spark validations. - Faster execution cadence: in the same month, “vision + voice” were both open-sourced; the gap from “releasing weights” to “giving a reproducible deploy” is now just one step — the key crossing that turns a model into productive capacity.
Actionable conclusion: if you track the domestic LLM / embodied chain, Step’s moves this month are the least ambiguous execution signal — dual-line open source + dual-engine deploy readiness beats any “coming soon.” Directly relevant to my own OpenInfer / Qwen3-4B DFlash: Step-3.7-Flash has official draft paths on both vLLM and SGLang, showing that “parallel draft + domestic / low-card deploy” is already a reusable pattern.
Worth saying separately: last issue (0915) covered “auditable speculative decoding (dual-key watermark)”; this issue Step ships mature deploys on both vLLM and SGLang, effectively pushing the “domestic MoE speculative-decode stack” from demo to production. Read together, speculative decoding is now both “safe for production” (watermark) and “runnable on domestic cards” (Step deploy ready).
2. vLLM & SGLang Community Tracking
Version status: vLLM stable v0.29.0 (09-09), next candidate v0.29.1rc0 (09-13); SGLang v0.5.19 (09-05) remains the latest stable, no new tag in-window. This period’s increments come mainly from the main branch, with no formal release (same cadence as 0915).
vLLM (main · engineering maturity + multi-backend reinforcement)
New features / architecture evolution:
- Rust Core keeps maturing: added
--hf-overridesand a per-request preemption histogram — taking scheduler observability from “global aggregate” down to “single request,” making tail-latency localization sharper. - XPU / ROCm backend weekly reinforcement: continuing to spread v0.29.1rc0’s dual-key Gumbel-max watermark (#56122, separate draft/verify sampling keys for production compliance / anti-forgery) across more hardware.
- stable v0.29.0 already includes: EPD three-segment split (Encoder–Prefill–Decode), OffloadingConnector fix (no longer zero-hit under MTP/EAGLE), etc. (see 0915).
SGLang (v0.5.19 · default prefix cache + DFlash2 draft-net upgrade)
New features / major adaptations:
- Unified Radix Tree becomes the default prefix cache (BREAKING): previously opt-in, now default — re-check your hit rate, or the cached-benefit assumptions of old deployments will go stale.
- Beam Search goes request-level: incompatible with speculative decoding / PD disaggregation (pick one per request).
- DeepEP v2 fills multi-node MoE graph-captured Decode: the decode stage of cross-node MoE can now enter CUDA Graph too.
- DFlash2 draft-net upgrade: adds local convolution + candidate selector, Qwen3.8-27B single H200 at concurrency 1 hits 3.43×; MXFP4 W4A8 MoE + KDA fused accept further cuts verify overhead.
Production practice: DFlash2’s “local conv + candidate selector” is directionally clear — it aims to directly raise the acceptance length E[L] (see the deep-dive), rather than widening the draft tree. Main-branch increment, not yet in the v0.5.19 tag.
Under the Hood: DFlash2 Draft Net + Local Conv + Candidate Selector — Pushing “Acceptance Length E[L]” Up
Last issue (0915) we nailed down a formula: Speedup = E[L]×Ttarget / (Tdraft + Tverify), concluding “E[L] is the ceiling, don’t blindly widen the draft tree.” This issue DFlash2 shows a concrete way to raise that ceiling:
- Local convolution: injects n-gram / positional-locality awareness into the draft net so drafts fit the target model’s local patterns better → more draft tokens accepted (E[L]↑).
- Candidate selector: “generate + filter” inside the parallel draft tree, sending only high-confidence candidates to verification → higher pass rate for the same Tverify.
Qwen3.8-27B single H200 at concurrency 1 → 3.43× is the payoff of this route. It shares DNA with 0913’s “cache-boundary Pareto” and 0915’s “E[L] ceiling”: all hunt the optimal operating point on a tradeoff surface, only this issue the knob moves from “choose D” to “improve draft-net quality.”
My read: DFlash2’s direction proves it — the next phase of speculative-decode competition is not in “parallelism” but in “draft quality.” Whoever pushes E[L] higher and Tverify lower wins. For my own OpenInfer / Qwen3-4B DFlash, the next step is to empirically test whether a “local-conv-style draft net” can reproduce 3×+ on Qwen3-4B.
Standing Topics
PD disaggregation: vLLM main takes EPD three-segment split + 9 classes of KV connector (NIXL/Mooncake/LMCache) + llm-d (CNCF); SGLang side defaults the unified Radix Tree + DeepEP v2 multi-node graph capture. Crosswise: vLLM leans connector-pluggable, SGLang leans router + native unified-tree cache (#37506 unified-pool).
Architecture evolution: four speculative-decode routes (EAGLE-3 / MTP / DFlash / DSpark); this issue SGLang pushes DFlash to DFlash2 (local conv + candidate selector); sparse attention HiSparse, FP8/NVFP4 quantization (NVFP4 lm_head still a pit), multimodal diffusion keep advancing.
PyTorch vs transformers: no new “off transformers” PR this cycle. Standing position holds: pure-PyTorch in-house pays off only for flagship models; Step-3.7-Flash debugging needs transformers ≥ 5.0. Crosswise conclusion unchanged: in-house inference stacks suit high-frequency / flagship / extreme optimization, long-tail leans on transformers.
Step adaptation: StepAudio 3 (09-15, 98.9% realtime / 1.7% ASR-WER global #1) + Step-3.7-Flash deploy maturity — vLLM stepfun37 (FP8/BF16 MTP k=3; NVFP4 4 cards TP4 needs modelopt + FP8 KV cache + async-scheduling) / SGLang dev-step-3.7-flash (EAGLE multi-layer draft). Vision+voice dual-line open-sourced in the same month, cadence accelerating. Crosswise: vLLM first-class citizen > SGLang dev image.
3. AI Papers & Industry Hotspots
Today’s Focus (1 sentence)
Google DeepMind’s SigLIP 2 upgrades the vision encoder from “recognizes objects” to a “semantic + dense + multilingual” triple and has become the default vision tower of new-gen VLAs like Qwen3-VL — the very eyes that let robots go “from recognizing to doing”; while M-RoPE (Qwen3-VL) preserves image structure with 3D positional encoding and Expert Offload lets a DeepSeek V4 671B-class MoE run on limited VRAM, rounding out today’s “vision + inference landing” thread from the model layer down to the deploy layer — and on the industry side, Unitree’s IPO acceptance, Zhiyuan’s HK listing, UBTECH’s orders past ¥1.4B, XPeng’s IRON off the line, and the dual ADAS national standards are each putting a price tag on the embodied “see + act” stack.
Paper Core (SigLIP 2 · Google DeepMind, arXiv:2502.14786)
- One-line positioning: a four-loss unified recipe (sigmoid contrast + caption decoder + self-distill/mask-prediction + online debias) produces a semantic / dense / multilingual vision-language encoder, backward-compatible with SigLIP.
- Core idea: the caption decoder injects positional grounding → localization; self-distillation (last 20% of training, 50% patch-mask reconstruction) fills local semantic / dense; 90% English + 10% non-English debias → multilingual.
- Formula:
L = L_sig + λ1·L_cap + λ2·L_self;L_sigis batch-internal all-pair sigmoid binary classification. - Impact (numbers speak): RefCOCO 83.76% vs 64.05%, VOC segmentation 77.1 vs 72.0, NYUv2 depth 0.493 vs 0.576; now the 2025–26 default open-source VLA vision tower (Qwen3-VL uses SigLIP2-SO-400M). Embodied meaning = more accurate grasping / depth + overseas multilingual + NaFlex screen reading.
Operator Deep-Dive: M-RoPE (Qwen3-VL)
- Positioning: expands 1D RoPE into 3D (T / H / W) positional encoding, treating image patches / video frames as a structured grid.
- In use: Qwen3-VL (32B / 235B-A22B) Interleaved MRoPE (
mrope_section=[24,20,20], t/h/w interleaved high-low freq) + DeepStack + video timestamps; the vision tower is SigLIP-2. - Key points: zero extra params explicitly encode coordinates (vs 1D flattening loses structure); pit =
mrope_sectionmust match the weights, Qwen2-VL/2.5-VL/3get_rope_indexare not interchangeable, video T goes through the timestamp token.
Performance Optimization: Expert Offload
- Positioning: offload MoE cold experts to CPU / NVMe, keep hot experts resident on GPU + async prefetch, so a DeepSeek V4 671B-class MoE runs on limited VRAM.
- Representative work: MoE-Infinity (ICML 2023), DeepSpeed-MoE; orthogonal to EP and stackable.
- Benefit: GPU VRAM drops from E×size to K_hot×size (saves several ×), net latency near zero via overlap; embodied link = edge VLA “edge GPU + CPU co-inference” real-time.
Industry Hotspots (embodied companies / chain speed-dial · pinned)
- [Embodied] Unitree (688836 · Sci-Tech board IPO accepted / planning to raise ¥4.2B): raised $905M in August, G1 at ¥99k / B2 profitable 5 years running, but Goldman flags “not ready to execute functional tasks” → neutral-slightly-cautious; bullish on harmonic/actuator supply chain (688017 / 002050 / 601689) and robot ETF 562500.
- [Embodied] Zhiyuan Robotics (HK IPO target HK$40–50B; controls Shangwei 688585): 2025 shipped 5,168 units, global #1, 2026 target breaks 10,000 units → bullish on 688585 mapping.
- [Embodied] UBTECH (09880.HK, ~¥56.5B mcap, orders past ¥1.4B): Walker S2 locked $112M in orders, the only listed whole-machine vendor.
- [Embodied] XPeng (09868.HK): 2nd-gen VLA + first IRON off the production line on 9/8 — edge params up 3.5×, response 300% faster, 4D spatiotemporal + world model, spanning Robotaxi / IRON → Horizon Robotics 09660.HK benefits.
- [Regulation] Mandatory ADAS national standards: GB44721-2026 (effective 2027-7-1) + GB47955-2026 (L2 naming spec) → bullish on compliant ADAS supply chain; 2026 H1 humanoid financing $16.2B hit a record high, China accounting for 10 deals.
⚠️ Industry developments are not investment advice.
4. The One-Line Takeaway
Step’s “Step series” delivered its strongest execution combo this month — StepAudio 3 (98.9% realtime / 1.7% ASR-WER, global #1) tops Artificial Analysis and Step-3.7-Flash is deploy-ready on both vLLM/SGLang engines (NVFP4 4 cards) — a clean execution signal for the “Step chain”; SGLang v0.5.19 makes the unified Radix Tree the default prefix cache (re-check hit rate) and DFlash2’s draft net adds local conv + candidate selector to push the acceptance length E[L] up (Qwen3.8-27B single H200 concurrency 1 at 3.43×), sharing the same formula as 0915’s “E[L] ceiling” — moving the knob from “choose D” to “improve draft-net quality”; vLLM’s main keeps Rust Core mature + dual-key watermark steady. On the papers side, SigLIP 2 (DeepMind, arXiv:2502.14786) upgrades the vision tower to a semantic+dense+multilingual triple with a four-loss recipe and is now the default eyes of new-gen VLAs like Qwen3-VL, M-RoPE preserves structure with 3D positional encoding, Expert Offload lets a DeepSeek V4 671B MoE run on limited VRAM, and Unitree’s IPO acceptance, Zhiyuan’s HK listing, UBTECH’s orders past ¥1.4B, XPeng’s IRON off the line, and the dual ADAS national standards are each putting a price tag on the embodied “see + act” stack.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。