系列:每日AI热点

Daily AI Hotspot · 2026-09-01: SGLang Turns NVFP4 + FP8 lm_head Into Endless Repetition — And You Cannot Install the Fix on v0.5.18

A note on today’s sources: only the vLLM / SGLang community tracking digest is available today — there is no corresponding AI papers and industry hotspot daily. This entry therefore covers the inference-infrastructure side only, and the industry/papers section is omitted.

★ Most Worth Your Attention Today

SGLang’s NVFP4 model + FP8 lm_head bug: output repeats endlessly, and the fix is not installable.

This is the only item today that will outright destroy your production output — and the trouble is not the bug itself, it is the reachability of the fix.

Three layers of fact:

  1. Symptom: NVFP4 checkpoints with an FP8 lm_head produce endlessly repeating output on SGLang. Not mild degradation — obvious enough that your users notice it immediately.
  2. Fix: merged in main PR #35228.
  3. The problem: the v0.5.18 branch predates that fix, and sglang[all]==0.5.18 is pinned on PyPI — there is no normal pip installation path that gets you a version carrying the fix.

Actionable conclusion: if you need to run FP8-lm_head checkpoints, either switch to vLLM or build SGLang from main. There is no third option.

Worth saying separately: a situation where “the fix is in main but the release channel is pinned” deserves more caution than the bug itself. It means your dependency-management strategy must include an emergency path for building from main — not because you want to chase the bleeding edge, but because you need an escape hatch between a fix and the next tag.

2. vLLM & SGLang Community Tracking

There is no AI papers / industry source for this period (see the note at the top); what follows is the inference-infrastructure portion.

Version status: no new releases in the last 72 hours. vLLM v0.28.0 (8/26, 584 commits / 270 contributors); SGLang v0.5.18 (8/22, 710 PRs).

vLLM v0.28.0

New features:

Breaking changes:

ChangePRImpact
bitsandbytes moved out of core → external plugin#43529The quantization path needs a separate install
Transformers upgraded to 5.15.0—Custom model definitions will very likely need edits
reasoning_content output removed—Downstream consumers parsing that field must adapt

⚠️ Small-VRAM users, read this before upgrading: go back over the 7 breaking lines in the release notes and re-measure your capacity. This release moves both the quantization dependency (bitsandbytes externalized) and the default batch size (see below), and both change your peak memory.

Default change: max_num_batched_tokens 8192 → 16384. Single-GPU / small-VRAM setups can OOM.

SGLang v0.5.18

Performance:

Known pit (see the ★ section): the FP8 lm_head bug on NVFP4 models (output repeats endlessly); fixed in main PR #35228, but the v0.5.18 branch predates the fix and sglang[all]==0.5.18 is pinned on PyPI and cannot be installed → for FP8-lm_head checkpoints, use vLLM or build from main.

Under the Hood: vLLM 0.28’s Tiered KV Cache Sinking to Disk

The most worthwhile technical read of the day.

The path: GPU → CPU RAM → local disk.

The significance goes well beyond “one more storage tier.” It means:

The inference engine is turning from “a runtime that manages VRAM” into “a runtime that schedules model state across compute, VRAM, main memory, and disk.”

The direct benefit: long-context agent workloads no longer need to occupy the most expensive memory permanently for every cold session. An agent session that sits idle for half an hour can have its KV resting quietly on disk, leaving HBM for requests that are actually running.

The cost is equally clear: cold restore pays an I/O tail latency. What used to be read straight from HBM may now wait on a disk copy-back. That latency does not touch steady-state throughput, but it lands squarely on P99 — which is exactly where interactive experience is most sensitive.

Your operational metrics have to expand in step. “KV cache utilization” used to be enough; now you need at least three more:

  1. Per-tier hit rate (how much is served from GPU / CPU / disk respectively)
  2. Inter-tier migration bytes (how much data is shuttling between tiers)
  3. Restore latency distribution (how long a cold KV takes to come back from disk)

My read: this is the right roadmap, but it pushes inference-engine complexity to an entirely new order of magnitude — it is starting to look like an operating system. Tiering, swapping, hit rates, tail latency: these words used to belong to storage systems, and now they belong to an inference engine. If nobody on your team owns those metrics yet, now is the time to name someone.

Standing Topics

PD disaggregation: vLLM’s Model Runner V2 pushes E/P/D disaggregation plus tiered KV to disk; on the SGLang side it is DCP decode-context parallelism / FlashInfer all-to-all MoE routing / DeepSeek-V4 FlashMLA sparse prefill on by default.

Architecture evolution: three speculative decoding roadmaps (EAGLE / MTP / adaptive budget — vLLM’s adaptive budget improves DSpark TTFT by ~60%), sparse MLA, FP8/INT4/KV quantization, and multimodality.

Pure PyTorch vs transformers: no new “off transformers, in-house stack” PRs this cycle, but vLLM’s move to Transformers 5.15.0 plus externalized bitsandbytes is an unmistakable dependency-decoupling signal. The trend is toward minimizing hard dependencies and making custom operators easier.

Step adaptation:

The Step-Audio item deserves attention: RTF 0.0053 means roughly 189× real time. The fact that a streaming task like ASR can harvest MTP’s gains shows speculative decoding spreading beyond “text generation” into autoregressive workloads generally. One hour of audio processed in 19 seconds is not really an “acceleration” anymore — it changes whether this class of task belongs in an online pipeline at all.

3. The One-Line Takeaway

On a day with no new release, the two most valuable pieces of information pull in opposite directions: on the positive side, vLLM 0.28 sinks KV to disk, making an inference engine schedule state in tiers like an operating system for the first time; on the negative side, SGLang v0.5.18’s NVFP4 + FP8 lm_head turns output into endless repetition, and the fix is locked outside the release channel — go confirm right now that your dependency management has an escape hatch for building from main.


📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。