系列:每日AI热点

Daily AI Hotspot · 2026-08-18: TensorCast Turns KV Migration Into a Shared Layer, Cutting Agent Median TTFT by Up to 93.2%

Today’s keyword is decoupling. On the engineering side, a cross-framework infrastructure proposal appeared: TensorCast, from Peking University and StepFun, pulls KV migration and weight materialization out of the inference engine itself into a shared tensor-as-a-service layer, integrated with both vLLM and SGLang. On the research side, Meta’s V-JEPA 2 puts “self-supervised world model + zero-shot robot planning” onto a real robot arm — and tomorrow Unitree lists on the STAR Market.

★ Most Worth Your Attention Today

TensorCast (Peking University x StepFun, arXiv:2608.06007) — decoupling KV migration and weight materialization into a shared TaaS layer, with median TTFT down as much as 93.2% in high-concurrency multi-turn agent workloads.

Why single this out: over the past two years, KV cache tiering, offloading, and cross-instance reuse have been rebuilt independently by nearly every inference framework. vLLM has NIXL and Mooncake connectors, SGLang has HiCache and a Mooncake transfer engine, LMCache is a third stack entirely. The result is the same capability implemented three times, with mutually incompatible KV formats — cross-framework co-deployment is simply not on the table.

TensorCast pushes the capability down a layer: KV migration stops being the job of connectors inside each engine and becomes an independent tensor-as-a-service that any engine can call. It ships integrations for both vLLM and SGLang, which is a first for this community.

How to read that −93.2%: the gain concentrates in one specific combination, “high concurrency + multi-turn agents.” Every turn of a multi-turn conversation replays a very long prefix; once concurrency rises, replay traffic saturates prefill compute and TTFT explodes. Pulling prefix KV back from a shared layer eliminates that entire recomputation. The flip side: single-turn, short-context, low-concurrency workloads will not see this gain — don’t force-fit the number onto your own traffic.

Caveat: this figure comes from the paper and press coverage, not from a community CI reproduction. But the architectural direction is what matters: once the KV cache stops being an engine-internal data structure and becomes infrastructure addressable across frameworks, competition between inference engines shifts from “whose cache is faster” to “whose interface is cleaner.”

1. AI Industry & Paper Highlights

Paper: V-JEPA 2 (Meta FAIR, 1.2B, arXiv:2506.09985)

Operator: HCA (Heavily Compressed Attention)

HCA is the directory layer in DeepSeek V4’s three-tier “far-coarse / mid-retrieval / near-full” memory design: it fuses 128 tokens into a single directory entry (V4 Appendix B’s Manifold-Aware Projection, an invertible fusion), then runs dense attention over roughly 8K directory entries.

PropertyValue / conclusion
Compression ratio128:1, roughly 128x less compute
Must be denseA dense kernel over 8K KV entries is more regular and faster — it cannot be made sparse
Relation to CSAComplementary: CSA retrieves precisely, HCA surveys globally
ShippingDeepSeek V4-Pro / V4-Flash; not used by Qwen3.8-Max (which leans on SWA+GQA); MiniMax-M1/H3 use Lightning Attention

The link to embodied AI is direct: VLAs must handle long video history and action logs simultaneously, which is exactly what a “global survey + precise retrieval” structure like HCA is built for.

Performance: Sequence Parallelism + Ring Attention

A system-level approach to long-context training and inference that is mathematically exactly equivalent to full attention, while per-device activation memory becomes independent of sequence length. Representative work: Liu et al., ICLR 2024 (100M tokens on 64 devices), DeepSpeed Ulysses (all-to-all, 2.5x faster), and OpenRLHF / ring-flash-attention (production-grade). Both DeepSeek V4 and Qwen3.8-Max rely on this path for 1M contexts. Two practical details: communication overhead can be hidden behind matmuls, and Zig-Zag partitioning fixes the idle-tail-device problem caused by causal masking.

Industry roundup

2. vLLM & SGLang Community Tracking

Both frameworks are in a “high commit volume, no new tag” phase: vLLM stable is v0.27.1 (08-11, 79 commits), SGLang is v0.5.17 (08-08, 100 commits).

vLLM

SGLang

Deep dive: where vLLM #49793’s win actually comes from

Every MTP speculative decoding step ends with an all-reduce that synchronizes results across TP ranks, and only then generates draft tokens. Those were two sequential passes with a device-to-host copy sandwiched between them (so the host could take the argmax for drafts).

The PR fuses the all-reduce with draft-token generation and moves the draft argmax onto the GPU, deleting that D2H copy outright.

The interesting part is why it only shows up at concurrency 64 (+13.6%) and is just ±3% at low concurrency: at low concurrency the synchronization is already absorbed by GPU idle time and never becomes a bubble. As concurrency rises, this fixed overhead stacks across every request and turns into an observable bubble on the critical path. Fusing it fills the bubble.

Methodological value: normalizing by acceptance length shows the gain genuinely comes from eliminating communication rather than from “verifying a few more tokens.” That normalize-then-attribute step is a cheap way to check whether a performance PR actually fixed the right thing.

Standing topic: Step-series support

This cycle brought Step crash-fix PRs from outside contributors (vLLM #52115 / SGLang #35206), while StepFun’s own repos remain frozen at 06-01 / 04-03. Note the irony: today’s ★, TensorCast, is Peking University × StepFun work — academic output is flowing while upstream engineering merges stall; Step’s investment is clearly not on the open-source inference stack side.

3. The One-Line Takeaway

The most valuable thing today is not a percentage but an architectural signal: TensorCast wants to lift KV migration from “each framework’s private connector” to “a layer shared across frameworks” — if that path works, vLLM and SGLang converge at the cache layer and competition returns to schedulers and kernels; and until then, the two gains you can bank immediately are vLLM #49793 (+13.6% at concurrency 64) and SGLang #35070’s production A/B (+17.98% TPS/User).


Sources: vLLM PRs #52188 / #49793 / #50493 / #52084 / #43107 / #51538 / #52115; SGLang PRs #35070 / #34801 / #35022 / #34696 / #35206; TensorCast arXiv:2608.06007 and coverage from news.qq.com / pith.science.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。