0. Why P900 + HiCache Forces You to Look at NUMA / PCIe / NIC
HiCache puts the KV cache in host DRAM, treating it as one large host-side KV pool. But when a request hits host KV, that data is not consumed on the P900 right away — it usually has to travel one more leg:
Host KV (host DRAM)
|
| H2D (host to device)
v
PCIe
|
v
P900
|
v
HBM (device memory)
In other words, a high host KV hit rate does not guarantee high performance, because the PCIe transfer cost comes after the hit. If PCIe bandwidth is insufficient, latency is high, or the KV lives across a NUMA boundary, you can end up with:
Host KV hit rate up
|
v
PCIe transfer
|
v
TTFT up / TPS down
So in this scenario PCIe, NUMA and NIC must be read as one path. First a topology diagram to make the hardware relationships concrete, then a five-layer optimization diagram that turns “how to tune it” into actionable steps.
1. Topology: What Actually Happens Inside One Server
Figure 1: P900 + HiCache server topology. Green = the P900 accesses host KV in its own NUMA node (the ideal path); red dashed = the P900 crosses the NUMA interconnect to reach another node's host KV (one extra hop, higher latency, lower effective bandwidth).
The division of labour, in one line each:
- PCIe — the data channel between the P900 and the host / CPU (DMA, bandwidth, latency, plus the control plane: device discovery, MMIO, BAR, P2P).
- NUMA — which CPU memory node the host KV sits in, and whether the P900 sits close enough to it.
- NIC — for PD disaggregation / Mooncake / RDMA / KV transfer, the NIC also needs to be on the same NUMA node as the P900 and the host KV.
The core data path, one-line version:
P900 HBM -> PCIe -> PCIe Root -> NUMA Node -> Host DRAM -> Host KV
2. Five-Layer Optimization: From Topology to Actual Pinning
Figure 2: Five layers of NUMA optimization. From "see the topology" through pinning CPUs, KV and NIC, to the final step of not over-pinning: find the balance between locality and load.
3. What You Should Measure Is Not “NUMA On / Off”
The valuable experiment compares different locality configurations, not a simple on/off switch:
| Configuration | Purpose |
|---|---|
| Correct topology binding | See the upper bound of the locality win |
| Automatic NUMA (system assigned) | Baseline for comparison |
| Wrong / remote binding | Measure how large the NUMA penalty is |
| Shared KV + NUMA ON | Best-case locality for shared KV |
| Shared KV + NUMA OFF | How sensitive shared KV is to NUMA |
If the results come out as:
correct NUMA > automatic NUMA > remote NUMA
that is essentially proof that HiCache performance is clearly affected by NUMA locality — which is exactly why parameters like --numa-node 0 0 2 2 exist.
4. One-Line Summary
PCIe answers “how does the P900 exchange data with the host / CPU”; NUMA answers “which CPU memory node holds that host data and is it close enough to the P900”; and in PD / Mooncake setups the NIC also has to be nearby.
So what you are ultimately optimizing is an entire path:
P900 <-> PCIe Root Complex <-> NUMA Node <-> Host DRAM <-> Host KV
not a single --numa-node flag.
💬 留言
NaphJohn/LLM-blog尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。