系列:Inference Systems Infrastructure Notes

Tuning SGLang HiCache L2 on Kunlunxin P800/P900: Layout, DMA, NUMA and TP-Shared KV

0. The one-line main thread

Whether tiered caching pays off depends not only on how high the hit rate can go, but also on how low the cross-tier transfer cost can be pushed.

HiCache extends prefix KV from HBM (L1) to Host DRAM (L2) and then to external storage (L3), trading a slower but far larger store for prefix reuse. But on Kunlunxin this mechanism defaults to a net loss — whenever the extra cross-tier transfer overhead outweighs the compute saved by hits, the whole thing slows down. Building on the NUMA x PCIe x NIC topology from sys1, this post breaks down the full path of tuning L2 from working to fast on P800/P900.

1. Why tiered prefix caching is needed

1.1 The problem: L1 capacity vs prefix reuse

The Prefill stage of LLM inference repeats a lot of computation. In Agent and multi-turn traffic, system prompts, tool definitions, and history overlap heavily across requests, so prefix duplication is high. Radix Cache reuses that computation: keep computed prefix KV in HBM, skip recompute on the next hit.

But Radix Cache lives on XPU HBM, its capacity hard-capped by device memory — under high concurrency the cache is evicted within minutes. Device memory cannot grow, yet two idle tiers sit unused: hundreds of GB to TB of Host DRAM per node, and cluster-level external storage. HiCache is about extending prefix caching into those two tiers.

1.2 The tiered design

HiCache splits the prefix index from the KV data location: the Radix tree only answers which node a token sequence maps to; where that node’s KV actually lives is a separate concern.

HiCache three-tier prefix cache architecture HiCache Three-Tier Prefix Cache Radix tree indexes, KV data placed per tier L1 · XPU HBM 10 GB · direct hit · no overhead L2 · Host DRAM 100 GB – TB · PCIe (D2H / H2D) L3 · External Storage TB – PB · net+IO · via L2 evict down load up evict load net benefit = compute saved by hit-rate − cross-tier transfer hit-rate up and transfer cost down together decide the sign

Fig 1: L1/L2/L3. When a node is evicted from L1, its KV can sink to L2 instead of being dropped; on the next prefix hit it loads back to L1. L3 is the next-level backstop for L2 and can be shared across instances.

After HiCache, the Radix tree is unchanged: a token sequence still maps to one node, prefix match / refcount / eviction logic stay the same. Only where the node’s KV lives changes. That defines a metric running through this whole post:

net benefit = compute time saved by higher hit rate − extra cost of cross-tier data movement

Every later change either raises the hit rate (bigger L2 / better L2 efficiency / L3) or lowers cross-tier transfer cost — which splits into four parts: 1) transfer bandwidth, 2) layout-transpose overhead, 3) operator-launch overhead, 4) interference of extra transfers with the main inference loop.

1.3 The adaptation surface on Kunlunxin

Landing HiCache on P800/P900 touches three layers:

2. L2 adaptation and tuning

2.1 Layout choice: page_first_direct to layer_first

HiCache L2 has three KV layouts: layer_first, page_first, page_first_direct. The first build chose page_first_direct — whole page contiguous, single layer page_size tokens contiguous, suited to direct (DMA) copy.

A quick test was discouraging: with HiCache off, the first request took 5.95s; with it on, 15.44s, and even the all-hit Replay request got slightly worse. The cost tracks the amount of data written back, not a fixed one-time overhead.

The real problem is the DMA submission count. page_first_direct dims are (page_num, layer_num, page_size, 1, kv_dim), page outermost. On the host side, different layers within one page are adjacent and can merge, but the same layer across different pages is always layer_num × page_size apart and can never merge; on the device side, different layers are different buffers and inherently non-contiguous. Combined, the segment count is fixed at page_num × layer_num.

Switching to layer_first (dims (layer, size, 1, kv_dim)), same-layer pages become contiguous and merge into one segment. At 32K tokens, page_size 64, 78 layers: 512 pages × 78 = 39,936 segments; with layer_first the ideal is 78 contiguous copies — roughly two orders of magnitude fewer. Each segment needs its own address computation and a DMA descriptor; the count itself decides whether single-layer transfer reaches near-sequential bandwidth.

KV layout vs DMA segment count KV Layout Decides DMA Segment Count eg: 32K tokens · page_size 64 · 78 layers → 512 pages page_first_direct page outer dim · cross-layer within page each page split into layer_num segments segments = page_num × layer_num 512 × 78 = 39,936 ↓ ~2 orders of magnitude layer_first layer outer dim · pages merged same layer pages contiguous merged segments = layer_num 78 segments both sides same layout → no transpose → net benefit turns positive (cold 10.7s → 7.9s)

Fig 2: Layout choice decides the DMA segment count. layer_first makes host and device layouts identical, so no transpose is needed inside the operator; segments drop from page_num×layer_num to layer_num.

The main route is set: layer_first + bidirectional async. With host layout switched to layer_first, the copy operators become same-layout transfer_kv_per_layer_mla (H2D per layer) / transfer_kv_all_layer_mla (D2H whole); cold-start first request drops from 10.74s to 7.91s.

2.2 Move transfer submission out of the scheduler thread

In the first build, the CPU-side submission (Kernel Launch) of H2D/D2H copy operators was inlined in the scheduler thread’s cache_unfinished_req path. For D2H, a single transfer_kv_all_layer_direct_lf_pf submission typically took 1.3s, and the scheduler loop could not continue until it returned. One request writes back multiple chunks, enough to explain the extra 9 seconds.

The fix: give D2H and H2D each a daemon thread consuming a queue. start_writing() no longer inlines submission but enqueues the request; a backup thread submits on a dedicated stream. start_loading() does the same into a load queue. Async removes the scheduler blockage, but the non-hit path was still ~26% slower than no HiCache — so the cost is not only about who submits.

2.3 The real problem: DMA submission count

Async moves submission off the main thread, but the amount of data moved per writeback does not shrink. The cost is in the DMA submission count (see 2.1). Switching to layer_first cuts segments by ~two orders of magnitude, so single-layer transfer can approach sequential bandwidth. This is what flips the L2 main route from net loss to net gain.

2.4 The physical form of Host memory

Async fixes who submits, but two remaining costs come from the physical form of Host memory:

Measured: registered memory size is itself a cost — 80GB to 15GB, degradation falls monotonically from +20.6% to +8.8%; binding Host KV to the local NUMA per XPU topology removes most degradation (+11.5% to -2.1%).

2.5 Platform-side PCIe ordering: IDO

Hygon has a hardware constraint: of 128 high-speed PCIe lanes, 64 per CPU are fixed for CPU interconnect, leaving only 4 PCIe x16. If each P800 took a dedicated x16, the main NIC would have no PCIe left — so two XPUs share one PCIe x16 uplink. The direct consequence: binding strictly by P800 NUMA topology uses only half the machine memory (a typical 1.5T box gives ~750GB; split across 8 cards minus component overhead, ~75GB per card max), becoming a capacity bottleneck.

Forcing all NUMA nodes (e.g. H3C topology 1 1 3 3 5 5 7 7, SGLang set to 0 1 2 3 4 5 6 7) uses all memory but D2H bandwidth collapses. Root cause: the two XPUs under one Switch should write memory independently, but strict PCIe ordering makes their writes block each other.

NUMA topology and PCIe ordering NUMA Topology and PCIe Ordering (Hygon P800) two XPUs share one Switch x16 uplink PCIe Switch (SW0) XPU0 → NUMA0 D2H write txn ID=A XPU1 → NUMA1 D2H write txn ID=B NIC1 RDMA contention strict ordering: ID=A and ID=B block → D2H bandwidth down use IDO (ID-Based Ordering): different-source IDs no longer wait → bandwidth back

Fig 3: Two XPUs share one Switch x16 uplink. Strict PCIe ordering blocks the two D2H writes against each other; switching to ID-Based Ordering (IDO) lets different-source-ID writes proceed independently and D2H bandwidth recovers.

The fix is to flip the PCIe RC communication policy to ID-Based Ordering (IDO) via a CPU register, so writes from different source IDs no longer wait for each other. Applies to NUMA one-to-one binding scenarios; after enabling, bandwidth recovers as expected.

2.6 Write-policy semantic fix + Ghost List

L1 to L2 write policy defaults to write_back / write_through / write_through_selective (backup only when L1 hit count ≥ 2, filtering one-shot prefixes). But under PD-disaggregated Prefill, the same request is inserted into the Radix tree twice (cache_unfinished_req once, cache_finished_req once), so a single request pushes hit_count to 2, and write_through_selective silently degrades to write_through — a semantic surprise.

Two fixes:

  1. Suppress same-request double counting: skip hit_count accumulation at the finished stage, so hit_count again means distinct requests.
  2. Ghost List: the fix introduces a new problem — if the reuse gap exceeds the prefix’s L1 residence, the second arrival finds the node already evicted and hit_count reset, dropping L2 hit rate. Ghost List keeps an extra bounded Hash Map of page hashes evicted from L1 (page granularity, not node, because chunk boundaries and partial-hit lengths differ across requests; the page-boundary hash accumulates from root and encodes both content and position). On rebuild, scan backward; the first hit is the longest reuse prefix; write only the proven-reused segment. Key is 64-bit, OrderedDict gives O(1) add/lookup, ~120B per entry.
Write policy / Ghost List / TP shared 1 · write_through_selective semantic fix PD disagg inserts same req twice → hit_count inflated → silently degrades to write_through fix: skip accumulation at finished stage → hit_count again counts distinct requests 2 · Ghost List page-granularity hash (64-bit) · OrderedDict LRU · O(1) add/lookup · ~120B each scan backward, first hit = longest reuse prefix · write only proven segment · bounded 3 · TP group shared Host KV MAP_SHARED buffer · only tp0 does D2H · all ranks read same indices MLA/NSA replicated latent bit-identical → L2 effective cap ×8 · D2H traffic ÷8

Fig 4: Three software optimizations. Write-policy semantic fix plus Ghost List cut useless writebacks; TP group shared Host KV merges N redundant copies into one, enlarging effective capacity and shrinking D2H traffic.

This policy is off by default — its benefit is conditional: whether TTFT saved by fewer D2H writes beats TTFT lost to lower hit rate depends on L2 capacity and the reuse-distance distribution. The tighter L2 is, the more filtering pays (it keeps one-shot prefixes out so limited L2 holds only proven-reusable high-value KV).

2.7 TP group shared Host KV

This is the single biggest win. For MLA/NSA, the Host-side KV is a replicated latent — every TP rank holds the bit-identical same data (even with dp attention, redundancy remains when TP≠1). So TP=8 means: 8 identical Host memory copies (L2 effective capacity only 1/8 of nominal) + 8x D2H traffic (while PCIe bandwidth, IOMMU/TLB budget, and kernel-launch quota are per-device or host-shared).

Mechanism: one TP group shares a MAP_SHARED Host buffer; only rank 0 in the group does D2H; all ranks read with the same indices. Media constraints dictate /dev/shm + MADV_HUGEPAGE (XPU driver cannot register hugetlbfs, but THP-backed tmpfs registers fine), with MPOL_INTERLEAVE temporarily set across all online nodes during registration. A side effect: under shared mode, NUMA binding becomes optional (L2 memory is already interleaved across nodes). Correctness relies on an implicit invariant: host pool alloc/free must be SPMD-deterministic across ranks, with an optional self-check for cross-rank hash consistency.

3. Results

End-to-end test on GLM5 Workbuddy 24K–64K, prefix reuse rate 0.82, PD disagg, concurrency 10:

StageKey changeTTFTPrefix hit ratevs Base
BaseHiCache off, L1 Radix only28.83s46.28%—
Main routelayer_first + kernel backend + bidirectional async22.08s69.23%-23.42%
+ Host mem formNUMA / hugepage / LIFO20.86s69.23%-27.66%
+ TP shared Host KVdrop N copies, L2 to 400GB19.12s72.79%-33.69%
+ L2 capacity tuneback to 200GB, hit-rate saturates18.45s71.94%-36.01%
End-to-end TTFT waterfall End-to-End TTFT Waterfall GLM5 Workbuddy 24K–64K · reuse rate 0.82 · concurrency 10 Base · L1 Radix only 28.83s + layer_first + bidirectional async 22.08s · -23.4% + Host mem form (NUMA/hugepage/LIFO) 20.86s · -27.7% + TP group shared Host KV 19.12s · -33.7% + L2 capacity tune (200GB) 18.45s · -36.0% + cache-aware + dp-aware routing 14.47s · -42.6% 0 hit-rate 46% to 72% is the main gain; the 12.6pt from -23% to -36% is all transfer-path optimization

Fig 5: End-to-end TTFT waterfall. From 28.83s down to 18.45s; with cache-aware + dp-aware routing, TTFT drops further to 14.47s.

Two key conclusions:

The all-hit fit: T(n) = 779.1ms + 13.539us · n, where 779ms ≈ 735ms (shared deployment floor) + 44ms (L2-path-only). Measured H2D bandwidth ~7.40 GB/s. Break-even is around 3K tokens: only hits above 3K tokens pay off.

4. Summary and outlook

HiCache L2 on Kunlunxin P800/P900 is already practically useful, with an implementation centered on layout optimization, async transfer, NUMA affinity, and TP group sharing. But every number here is measured in a specific environment and should not be extrapolated to other platforms or workloads — when switching, re-attribute in the same order: confirm the host side (ordering, NUMA, huge pages) leaves no obvious single-transfer cost, then probe hit-rate marginal gain tier by tier, and only then touch dataset-sensitive switches like write policy.

The future is L3: Host DRAM alone is not enough for longer context and higher efficiency. The fastest, cheapest path is to pool the D node Host DRAM via Mooncake in PD disagg — D nodes run Decode and their Host DRAM is long underused; pooling needs no new hardware and gives Prefill instances sizable L3 capacity, still DRAM so cross-tier cost is far below disk; it also turns single-node private cache into cluster-shared cache, freeing hit rate from single-instance locality. Remaining questions: RDMA interference with L2 writeback, and pool-metadata consistency at scale.

5. Series navigation

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。