A note on today’s sources: only the vLLM / SGLang community tracking digest is available today — there is no corresponding AI papers and industry hotspot daily. This entry therefore covers the inference-infrastructure side only, and the industry/papers section is omitted.
★ Most Worth Your Attention Today
Agentic workloads have officially become the main battleground for vLLM / SGLang optimization — the win condition has shifted from “fixed 8k prompt throughput” to “whether a multi-turn agent session can still hit its history KV on the next round.”
This is the one item that should go straight into your architecture review — because it changes the criteria you use for capacity planning and engine selection.
Three layers of fact:
- Criterion migration: SemiAnalysis’s deep comparison notes that both projects have rewritten KV cache storage, retrieval, eviction, and cross-node migration around hybrid attention (Full + Sliding-Window + linear/recurrent state). The metric moved from “fixed 8k prompt throughput” to “whether a multi-turn session can still hit its history on the next round.”
- Three classes of state, completely different cost structures: Full KV is expensive but reusable; the SWA tail is cheap and rebuildable; recurrent state is cheap but not rebuildable — “how much to offload / who to evict” has degraded from a unified policy into three tiered policies.
- The pitfall to fear most (from this period’s correctness fix): when ROCm ring-cache reuses a still-referenced slot, it yields wrong output, not slow output — silent errors are more dangerous than slowness because they never raise an alarm.
Actionable conclusion: if you are building an inference substrate for agents / multi-turn dialogue, change your evaluation metric from “first-packet throughput” to “cross-turn KV hit rate + state-migration cost,” and set separate policies for the Full / SWA / recurrent state classes; on ROCm deployments, confirm the ring-cache fix is merged.
Worth emphasizing: this is not a compute shortage — it is that “who should stay in HBM” has become a harder question. The variable-length + forking nature of agentic loads makes static assumptions like “specialize by length” or “evict by LRU” all fail. The inference engine is turning from “a runtime that manages VRAM” into “a runtime that schedules three classes of state across GPU / main memory / disk” — an extension of the same direction as the 9/1 KV-to-disk entry.
2. vLLM & SGLang Community Tracking
There is no AI papers / industry source for this period (see the note at the top); what follows is the inference-infrastructure portion.
Version status: no new releases in the last 72 hours. vLLM v0.28.0 (8/26); SGLang v0.5.18 (710 PRs / 212 contributors).
vLLM v0.28.0
The three most practical KV Offload changes (store path):
- Skip store if the same transfer is already in-flight (a concurrent shared prefix pays only once).
- Write only the incremental KV range (don’t rewrite the whole prefix every round).
- Store no longer depends on the block still being in HBM (scheduled work is not lost on eviction).
Load path: scheduling lookup went async (connector off the step critical path) + compact zero-copy lookup key + receiver-side parallel load + pre-built Mooncake key string.
Correctness ledger (hybrid state): emit cache events by hybrid cache group; EAGLE speculative state propagates across Mooncake merge groups / SimpleCPU coordinator.
Three ROCm changes: Kimi-K3 KDA writes directly into the layer output buffer (saves one device copy per layer per token); AITER sparse-MLA decode kernel → AgentX output throughput +5.22% with significantly lower ITL; pending highlight — DeepSeek V4 C4A selector moved to gfx950 hybrid AITER/native → end-to-end 1.21×–1.76×.
PD disaggregation: Kimi-K3 conv+ssm recurrent state travels with attention KV over MoRI-IO (1P1D); MI355X TP8 cross-node RDMA, both legs run DSpark speculative decoding.
SGLang v0.5.18
Sliding-window allocator: window / prefix pages share a pool, window is greedier → pages free immediately on leaving the window, compute-lock cap set to a single window, stale full-KV cleared. Legacy blind spot: a forked request still holds reusable full-KV but the fork-point window state was already released → whole-prefix recompute (work in progress: “retain SWA state at the fork point for fork inheritance”).
HiCache asymmetric offload: offload only the full-attention cache, rebuild the short SWA tail on return (move the expensive half, rebuild the cheap half faster); on AMD, staged write-back does not block the engine.
Recurrent-state gap closed: FlashInfer GDN checkpoints let recurrent state participate in prefix reuse for the first time — at 92.4% hit rate, 47,771 → 53,004 tok/s/GPU.
Variable-length context: context length becomes a runtime scalar (avoids per-request kernel compilation) → AgentX at concurrency 384, output throughput +26.75%, average TTFT −36.25% (not faster compute — just no compilation).
Correctness fixes: ROCm ring-cache reusing a still-referenced slot → wrong output (not slow output); also, a rank fed by chunked-prefill continue-compute always wins prefill-first → a configurable decode interval forces decode rounds to be inserted.
Under the Hood: Today’s Connecting Thread — Models Became “Hybrids”
String all of today’s changes together and they are one story:
- Models went from pure Full Attention to a Full + SWA + linear/recurrent state hybrid.
- The three state classes have completely different cost structures: Full KV is expensive but reusable, the SWA tail is cheap but rebuildable, recurrent state is cheap but not rebuildable.
- “Whether to offload / how much / who to evict” degraded from a unified policy into three tiered policies.
- The agentic load’s variable-length + forking nature makes static assumptions like “specialize by length” and “evict by LRU” all fail.
Implication for speculative decoding (directly relevant to your work): speculative state (EAGLE / MTP / DSpark) is now also “state that must migrate and be reused alongside KV.” The gain from speculative decoding is no longer just “accept a few more tokens per step” — it is strongly coupled with cache hit rate and state-migration cost. Profiling that only looks at acceptance rate misses an entire block of gain/loss. This pushes 9/2’s “adaptive budget at zero cost” one step further: the budget saves on redundancy, but only if the state-migration ledger is correct.
Two Hard Domestic News Items
- Zhipu GLM-5 / LayerSplit: the Coding-Agent load is “long context, high prefix hit rate” → Prefill-dominated. First use timeout Abort to suppress TTFT, HiCache to relieve Prefill-side KV capacity; LayerSplit lets each GPU under CP hold only part of the layer KV, the holding rank broadcasts that layer’s Cache before Attention and overlaps it with indexer computation — extra communication is only Indexer Cache broadcast ≈ 1/8 of the KV Cache (submitted as SGLang PR #22811).
- FlagOS three Day0s in four days (Qianwen / GLM / Hunyuan): MoE fused expert (net ~0.6ms) + dense/projection reusing the native fast path (~0.5ms), saving about 8.66ms per step; but operator speedup ≠ end-to-end speedup — you must bring metadata construction into the CUDA Graph (add/rewrite 3 metadata kernels). The contrast is compelling: native only +0.93%, FlagOS from partial to complete graph entry +7.46%. Also dug out an upstream precision defect (GDN Packed Decode’s beta FP32 result unnecessarily rounded back to BF16 → long-sequence recurrent-state error accumulation, PR #53877).
Standing Topics
| Topic | Today’s status |
|---|---|
| PD disaggregation | vLLM: Kimi-K3 recurrent state travels with KV via MoRI-IO (MI355X TP8 RDMA + both legs DSpark); SGLang: HiCache asymmetric offload + AMD staged write-back |
| Architecture evolution | Hybrid attention becomes the default assumption; GLM-5.3-Flash uses a plugin layer to absorb multi-chip differences; FlagOS’s same operator set carries into the Qwen4 model family |
| KV cache tiering | HBM → CPU DRAM (pinned + async DMA) → local NVMe → remote (Redis / Mooncake / InfiniStore); LMCache CacheBlend solves out-of-order RAG chunk reuse, TTFT ~2–3× improvement |
| Cache-affinity routing | New entry: Meta CacheRoute — high-frequency requests return to existing cache nodes, balancing cache locality against load balancing; the most easily overlooked hit-rate killer at scale |
| Step adaptation | No new additions this week; maintain Step-3.7-Flash NVFP4 four-card local deployment + MTP/EAGLE speculative scope |
Ops & Security
- Version floor: CVE-2026-54234 (DoS, fixed 0.24.0), CVE-2026-73559 (completion fan-out DoS, fixed 0.26.0) → production recommended ≥0.26.0, ideally 0.28.0.
- KV offload CLI:
--kv-offloading-size(GiB, sum across ranks when TP>1) +--kv-offloading-backend native|lmcache, activated only when size is set. - ReplaySSM:
--use-replayssmcaches recent SSM inputs, skips per-step full-state store, and writes back checkpoints on flush; requires a mamba backend and works standard decode only (no speculation). - Weight offload: OffloadConfig offers
uva(zero-copy) /prefetch(grouped async prefetch) /auto.
3. The One-Line Takeaway
On a day with no new release, the highest value comes from SemiAnalysis pointing both engines’ optimization battleground at agentic workloads — the win condition shifts from “fixed 8k throughput” to “whether a multi-turn session still hits its history KV,” and the engineering that backs it is “separate policies for the Full / SWA / recurrent state classes”; the things to fear most are that ROCm ring-cache reusing a still-referenced slot yields wrong output rather than slow output, and FlagOS’s reminder that “operator speedup ≠ end-to-end speedup” (metadata outside the CUDA Graph gets only +0.93%, complete graph entry gets +7.46%) — the inference engine is now scheduling state in tiers like an operating system, and your evaluation metrics and ops ledger must be upgraded to match.
📬 Want this kind of daily tracking in your inbox? Leave your email or join my list 👉 1023628035@qq.com
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。