Today’s keyword is decoupling. On the engineering side, a cross-framework infrastructure proposal appeared: TensorCast, from Peking University and StepFun, pulls KV migration and weight materialization out of the inference engine itself into a shared tensor-as-a-service layer, integrated with both vLLM and SGLang. On the research side, Meta’s V-JEPA 2 puts “self-supervised world model + zero-shot robot planning” onto a real robot arm — and tomorrow Unitree lists on the STAR Market.
★ Most Worth Your Attention Today
TensorCast (Peking University x StepFun, arXiv:2608.06007) — decoupling KV migration and weight materialization into a shared TaaS layer, with median TTFT down as much as 93.2% in high-concurrency multi-turn agent workloads.
Why single this out: over the past two years, KV cache tiering, offloading, and cross-instance reuse have been rebuilt independently by nearly every inference framework. vLLM has NIXL and Mooncake connectors, SGLang has HiCache and a Mooncake transfer engine, LMCache is a third stack entirely. The result is the same capability implemented three times, with mutually incompatible KV formats — cross-framework co-deployment is simply not on the table.
TensorCast pushes the capability down a layer: KV migration stops being the job of connectors inside each engine and becomes an independent tensor-as-a-service that any engine can call. It ships integrations for both vLLM and SGLang, which is a first for this community.
How to read that −93.2%: the gain concentrates in one specific combination, “high concurrency + multi-turn agents.” Every turn of a multi-turn conversation replays a very long prefix; once concurrency rises, replay traffic saturates prefill compute and TTFT explodes. Pulling prefix KV back from a shared layer eliminates that entire recomputation. The flip side: single-turn, short-context, low-concurrency workloads will not see this gain — don’t force-fit the number onto your own traffic.
Caveat: this figure comes from the paper and press coverage, not from a community CI reproduction. But the architectural direction is what matters: once the KV cache stops being an engine-internal data structure and becomes infrastructure addressable across frameworks, competition between inference engines shifts from “whose cache is faster” to “whose interface is cleaner.”
1. AI Industry & Paper Highlights
Paper: V-JEPA 2 (Meta FAIR, 1.2B, arXiv:2506.09985)
- One-line positioning: 1M hours of self-supervised video pretraining + 62 hours of Droid fine-tuning → 65–80% zero-shot pick-and-place success on a Franka arm.
- Four core ideas:
- No pixel reconstruction — it predicts in latent space only:
z_{t+1} = Predictor(z_{t-k..t}, a_t), with the loss being a D-dimensional L2. This is the sharpest divide from generative world models: it refuses to spend capacity modeling pixel detail irrelevant to decision-making. - An EMA teacher encoder supplies stable targets (momentum 0.99→0.999).
- Masked latent prediction (MAE-style), which saves compute and yields more robust representations.
- A planning head “imagines” K steps forward in latent space, and MPC selects the action sequence closest to the goal-image embedding.
- No pixel reconstruction — it predicts in latent space only:
- Why it matters: this is the first time a world model has gone from paper to physical hardware. Read alongside GR00T-Dreams (08-10) and Dreamer V3 (08-13), it completes a clean comparison: the self-supervised and RL branches of world models both reaching industrial viability.
Operator: HCA (Heavily Compressed Attention)
HCA is the directory layer in DeepSeek V4’s three-tier “far-coarse / mid-retrieval / near-full” memory design: it fuses 128 tokens into a single directory entry (V4 Appendix B’s Manifold-Aware Projection, an invertible fusion), then runs dense attention over roughly 8K directory entries.
| Property | Value / conclusion |
|---|---|
| Compression ratio | 128:1, roughly 128x less compute |
| Must be dense | A dense kernel over 8K KV entries is more regular and faster — it cannot be made sparse |
| Relation to CSA | Complementary: CSA retrieves precisely, HCA surveys globally |
| Shipping | DeepSeek V4-Pro / V4-Flash; not used by Qwen3.8-Max (which leans on SWA+GQA); MiniMax-M1/H3 use Lightning Attention |
The link to embodied AI is direct: VLAs must handle long video history and action logs simultaneously, which is exactly what a “global survey + precise retrieval” structure like HCA is built for.
Performance: Sequence Parallelism + Ring Attention
A system-level approach to long-context training and inference that is mathematically exactly equivalent to full attention, while per-device activation memory becomes independent of sequence length. Representative work: Liu et al., ICLR 2024 (100M tokens on 64 devices), DeepSpeed Ulysses (all-to-all, 2.5x faster), and OpenRLHF / ring-flash-attention (production-grade). Both DeepSeek V4 and Qwen3.8-Max rely on this path for 1M contexts. Two practical details: communication overhead can be hidden behind matmuls, and Zig-Zag partitioning fixes the idle-tail-device problem caused by causal masking.
Industry roundup
- Unitree Robotics lists on the STAR Market 8/19 (tomorrow): IPO price 150.80 CNY, 61B CNY valuation, 8,288x oversubscribed, with 2.022B CNY of the raise going to embodied foundation models — the first time this exceeds investment in the robot body itself. That is the most important line of the day: the company is betting on the brain, not the chassis.
- Mech-Mind passes the HKEX hearing: “the first stock spanning embodied eyes, brain, and hands,” with 2025 revenue of 389M CNY, 22.1% global share in 3D vision, and 64.6% gross margin.
- Infinite Origin closes ~1B CNY across A/A+ rounds: its causal world model AtomBrain is deployed in 100+ scenarios across 30+ cities, with cumulative orders of 500M CNY (a single order worth 260M) — resonating directly with today’s V-JEPA 2 theme.
- Peking University x BAAI x AgiBot: autonomous dexterous-hand serving (8/13): the first fully autonomous closed loop combining whole-body locomotion with fine dexterous manipulation, with 20 kHz pulsed visual response — 10x faster.
- Omdia data: H1 2026 China production 40,000 units, global shipments 19,100 (+272% YoY), Chinese vendors at 97%; Morgan Stanley raised its China humanoid shipment forecast to 50,000 units for the second time this year.
- Models: Qwen3.8-27B passes one million downloads in 48 hours, with 151,000 Qwen derivative models (2.6x Meta’s); DeepSeek V4-Pro GA on 8/13 with a leap in agent capability, 1M-token context, and output priced at $0.87 per million tokens.
2. vLLM & SGLang Community Tracking
Both frameworks are in a “high commit volume, no new tag” phase: vLLM stable is v0.27.1 (08-11, 79 commits), SGLang is v0.5.17 (08-08, 100 commits).
vLLM
- #52188: makes Kimi-K3’s DCP and DSpark coexist for the first time. They were previously mutually exclusive, forcing K3 users to choose between a parallelism strategy and speculative decoding.
- #49793: fuses the “per-step trailing all-reduce” of MTP with draft-token generation, taking the argmax locally for drafts and eliminating one D2H copy. Output throughput +13.6% at concurrency 64; ±3% at low concurrency.
- #50493: DCP partial prefix hits. #52084: DSV4 sparse top-k prefill. #43107: CI gate added to catch GPU↔CPU synchronization. #51538: DSV4 sparse MLA, three modes end to end.
SGLang
- #35070: production A/B of PD disaggregation measures +17.98% TPS/User. The most persuasive number of the cycle — PD disaggregation finally has production-grade evidence rather than benchmark-only claims.
- #34801: HiCache backs off to preserve decode-side KV. #34793: HiCache L2 flattening.
- #35022–35028: config-bag refactor consolidating
ServerArgsglobal state — a textbook “paying down maintainability debt” commit series. - #34696: DSpark gains logprobs support.
Deep dive: where vLLM #49793’s win actually comes from
Every MTP speculative decoding step ends with an all-reduce that synchronizes results across TP ranks, and only then generates draft tokens. Those were two sequential passes with a device-to-host copy sandwiched between them (so the host could take the argmax for drafts).
The PR fuses the all-reduce with draft-token generation and moves the draft argmax onto the GPU, deleting that D2H copy outright.
The interesting part is why it only shows up at concurrency 64 (+13.6%) and is just ±3% at low concurrency: at low concurrency the synchronization is already absorbed by GPU idle time and never becomes a bubble. As concurrency rises, this fixed overhead stacks across every request and turns into an observable bubble on the critical path. Fusing it fills the bubble.
Methodological value: normalizing by acceptance length shows the gain genuinely comes from eliminating communication rather than from “verifying a few more tokens.” That normalize-then-attribute step is a cheap way to check whether a performance PR actually fixed the right thing.
Standing topic: Step-series support
This cycle brought Step crash-fix PRs from outside contributors (vLLM #52115 / SGLang #35206), while StepFun’s own repos remain frozen at 06-01 / 04-03. Note the irony: today’s ★, TensorCast, is Peking University × StepFun work — academic output is flowing while upstream engineering merges stall; Step’s investment is clearly not on the open-source inference stack side.
3. The One-Line Takeaway
The most valuable thing today is not a percentage but an architectural signal: TensorCast wants to lift KV migration from “each framework’s private connector” to “a layer shared across frameworks” — if that path works, vLLM and SGLang converge at the cache layer and competition returns to schedulers and kernels; and until then, the two gains you can bank immediately are vLLM #49793 (+13.6% at concurrency 64) and SGLang #35070’s production A/B (+17.98% TPS/User).
Sources: vLLM PRs #52188 / #49793 / #50493 / #52084 / #43107 / #51538 / #52115; SGLang PRs #35070 / #34801 / #35022 / #34696 / #35206; TensorCast arXiv:2608.06007 and coverage from news.qq.com / pith.science.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。