系列:Inference Systems Infrastructure Notes

Inference Systems Infrastructure Notes (7): DCP — Decode Context Parallelism, sharding KV cache along the sequence dimension

0. The One-Line Thread

In long-context serving the Decode-phase bottleneck is not FLOPs but redundant KV-cache replication across GPUs. DCP (Decode Context Parallelism) shards the KV cache along the sequence dimension across GPUs and reuses a Scratch Buffer for collective comms — saving memory and cutting latency, with up to 1.5%~4% end-to-end throughput gain. This post links it with HiCache (sys1), HiSparse (sys6) and P/D disaggregation from this series.

1. Why DCP: the KV wall in Decode

Under standard Tensor Parallelism (TP), each layer places its KV cache on every GPU that participates in that layer. The longer the sequence, the larger the KV — but TP splits along the hidden dimension, not the sequence dimension, so:

Decode is token-by-token, compute-light but memory-bound, so “fits in memory, fetched fast” matters more than matmul throughput. DCP targets exactly this.

2. Core mechanism: KV sharded along the sequence dimension

DCP changes KV from “full copy per GPU” to “each GPU keeps only its slice”:

TP split:  KV replicated across hidden dim → every GPU holds full-sequence KV (redundant)
DCP split: KV sharded across sequence dim   → every GPU holds only [start:end] (no redundancy)

3. Communication optimization: reuse the Scratch Buffer

Sharding costs cross-GPU comms every Decode step. DCP’s key engineering trick is reusing a Scratch Buffer for all_gather / reduce_scatter:

4. Decode and prefill phases it touches

DCP changes more than Decode — it reaches preprocessing too:

5. Where it pays off, and what it synergizes with

DCP’s value is sharpest when:

6. Position among this series’ optimizations

OptimizationSolvesDimension
HiCache (sys1)which NUMA node hosts KV, how NIC joinsHW topology / data link
HiSparse (sys6)HBM as cache, host DRAM holds full KV, GPU keeps hot slotstiered cache (capacity wall)
DCP (sys7)KV sharded on sequence dim + Scratch Buffer commsparallel sharding (redundancy / comms wall)
P/D disaggregationPrefill and Decode instances decoupledarchitecture split

Three complementary angles: HiCache decides where to put, HiSparse decides how much to keep, DCP decides how to shard in parallel.

7. ⚠️ Disambiguation: three meanings of DCP

In high-performance inference frameworks (vLLM / SGLang), DCP almost always means Decode Context Parallelism. Elsewhere it can also mean:

Always confirm the context before writing/reading — don’t confuse “parallel sharding” with “prompt compression” or “visual pruning”.

8. Takeaway

DCP is the memory/comms optimization for long-context Decode: shard KV along the sequence dimension to kill redundant replication, reuse a Scratch Buffer to cut comms latency, and synergize natively with Prefix Caching, speculative decoding and P/D disaggregation. Together with HiSparse and HiCache it forms the three faces of long-context serving — capacity wall, data link, parallel sharding.

Note: compiled from the user-provided DCP technical brief on 2026-09-08. The 1.5%~4% throughput is a framework-measured range; real gains vary with sequence length, batch size and comms topology.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。