A note on today’s sources: only the vLLM / SGLang community tracking digest is available today — there is no corresponding AI papers and industry hotspot daily. This entry therefore covers the inference-infrastructure side only, and the industry/papers section is omitted.
★ Most Worth Your Attention Today
SGLang’s NVFP4 model + FP8 lm_head bug: output repeats endlessly, and the fix is not installable.
This is the only item today that will outright destroy your production output — and the trouble is not the bug itself, it is the reachability of the fix.
Three layers of fact:
- Symptom: NVFP4 checkpoints with an FP8 lm_head produce endlessly repeating output on SGLang. Not mild degradation — obvious enough that your users notice it immediately.
- Fix: merged in main PR #35228.
- The problem: the v0.5.18 branch predates that fix, and
sglang[all]==0.5.18is pinned on PyPI — there is no normal pip installation path that gets you a version carrying the fix.
Actionable conclusion: if you need to run FP8-lm_head checkpoints, either switch to vLLM or build SGLang from main. There is no third option.
Worth saying separately: a situation where “the fix is in main but the release channel is pinned” deserves more caution than the bug itself. It means your dependency-management strategy must include an emergency path for building from main — not because you want to chase the bleeding edge, but because you need an escape hatch between a fix and the next tag.
2. vLLM & SGLang Community Tracking
There is no AI papers / industry source for this period (see the note at the top); what follows is the inference-infrastructure portion.
Version status: no new releases in the last 72 hours. vLLM v0.28.0 (8/26, 584 commits / 270 contributors); SGLang v0.5.18 (8/22, 710 PRs).
vLLM v0.28.0
New features:
- DCP + fused FlashKDA + GEMM-RS
- Shared-expert sharding, saving ~17 GiB per GPU
- Adaptive speculative token budget, improving DSpark TTFT by ~60%
Breaking changes:
| Change | PR | Impact |
|---|---|---|
| bitsandbytes moved out of core → external plugin | #43529 | The quantization path needs a separate install |
| Transformers upgraded to 5.15.0 | — | Custom model definitions will very likely need edits |
reasoning_content output removed | — | Downstream consumers parsing that field must adapt |
⚠️ Small-VRAM users, read this before upgrading: go back over the 7 breaking lines in the release notes and re-measure your capacity. This release moves both the quantization dependency (bitsandbytes externalized) and the default batch size (see below), and both change your peak memory.
Default change: max_num_batched_tokens 8192 → 16384. Single-GPU / small-VRAM setups can OOM.
SGLang v0.5.18
Performance:
- Engine startup 2.38× faster (see the deep-dive below)
- Kimi K3 on AMD MI355X: throughput 1.37–1.77×
- NVFP4 checkpoints re-quantized online to MXFP4 run on AMD at 97.5–100% accuracy
Known pit (see the ★ section): the FP8 lm_head bug on NVFP4 models (output repeats endlessly); fixed in main PR #35228, but the v0.5.18 branch predates the fix and sglang[all]==0.5.18 is pinned on PyPI and cannot be installed → for FP8-lm_head checkpoints, use vLLM or build from main.
Under the Hood: vLLM 0.28’s Tiered KV Cache Sinking to Disk
The most worthwhile technical read of the day.
The path: GPU → CPU RAM → local disk.
The significance goes well beyond “one more storage tier.” It means:
The inference engine is turning from “a runtime that manages VRAM” into “a runtime that schedules model state across compute, VRAM, main memory, and disk.”
The direct benefit: long-context agent workloads no longer need to occupy the most expensive memory permanently for every cold session. An agent session that sits idle for half an hour can have its KV resting quietly on disk, leaving HBM for requests that are actually running.
The cost is equally clear: cold restore pays an I/O tail latency. What used to be read straight from HBM may now wait on a disk copy-back. That latency does not touch steady-state throughput, but it lands squarely on P99 — which is exactly where interactive experience is most sensitive.
Your operational metrics have to expand in step. “KV cache utilization” used to be enough; now you need at least three more:
- Per-tier hit rate (how much is served from GPU / CPU / disk respectively)
- Inter-tier migration bytes (how much data is shuttling between tiers)
- Restore latency distribution (how long a cold KV takes to come back from disk)
My read: this is the right roadmap, but it pushes inference-engine complexity to an entirely new order of magnitude — it is starting to look like an operating system. Tiering, swapping, hit rates, tail latency: these words used to belong to storage systems, and now they belong to an inference engine. If nobody on your team owns those metrics yet, now is the time to name someone.
Standing Topics
PD disaggregation: vLLM’s Model Runner V2 pushes E/P/D disaggregation plus tiered KV to disk; on the SGLang side it is DCP decode-context parallelism / FlashInfer all-to-all MoE routing / DeepSeek-V4 FlashMLA sparse prefill on by default.
Architecture evolution: three speculative decoding roadmaps (EAGLE / MTP / adaptive budget — vLLM’s adaptive budget improves DSpark TTFT by ~60%), sparse MLA, FP8/INT4/KV quantization, and multimodality.
Pure PyTorch vs transformers: no new “off transformers, in-house stack” PRs this cycle, but vLLM’s move to Transformers 5.15.0 plus externalized bitsandbytes is an unmistakable dependency-decoupling signal. The trend is toward minimizing hard dependencies and making custom operators easier.
Step adaptation:
- Step-3.7-Flash’s open weights have been validated in local deployments on Mac Studio M4Max / DGX Spark / AMD AI Max+395. → The edge and workstation path is open.
- Step-Audio 2.5 uses MTP to drive ASR-mode RTF down to 0.0053 — 1 hour of audio in 19 seconds.
- For specifics on vLLM / SGLang deployment, defer to the official StepFun repositories.
The Step-Audio item deserves attention: RTF 0.0053 means roughly 189× real time. The fact that a streaming task like ASR can harvest MTP’s gains shows speculative decoding spreading beyond “text generation” into autoregressive workloads generally. One hour of audio processed in 19 seconds is not really an “acceleration” anymore — it changes whether this class of task belongs in an online pipeline at all.
3. The One-Line Takeaway
On a day with no new release, the two most valuable pieces of information pull in opposite directions: on the positive side, vLLM 0.28 sinks KV to disk, making an inference engine schedule state in tiers like an operating system for the first time; on the negative side, SGLang v0.5.18’s NVFP4 + FP8 lm_head turns output into endless repetition, and the fix is locked outside the release channel — go confirm right now that your dependency management has an escape hatch for building from main.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。