系列:Inference Systems Infrastructure Notes

Inference Systems Infrastructure Notes (1): The NUMA x PCIe x NIC Data Path Under P900 + HiCache

0. Why P900 + HiCache Forces You to Look at NUMA / PCIe / NIC

HiCache puts the KV cache in host DRAM, treating it as one large host-side KV pool. But when a request hits host KV, that data is not consumed on the P900 right away — it usually has to travel one more leg:

Host KV (host DRAM)
     |
     |  H2D (host to device)
     v
   PCIe
     |
     v
   P900
     |
     v
   HBM (device memory)

In other words, a high host KV hit rate does not guarantee high performance, because the PCIe transfer cost comes after the hit. If PCIe bandwidth is insufficient, latency is high, or the KV lives across a NUMA boundary, you can end up with:

Host KV hit rate up
     |
     v
   PCIe transfer
     |
     v
TTFT up / TPS down

So in this scenario PCIe, NUMA and NIC must be read as one path. First a topology diagram to make the hardware relationships concrete, then a five-layer optimization diagram that turns “how to tune it” into actionable steps.

1. Topology: What Actually Happens Inside One Server

P900 + HiCache server topology: NUMA x PCIe x NIC NUMA Node 0 NUMA Node 2 PCIe Root Complex P900 HBM NIC RDMA Host DRAM Host KV (HiCache) PCIe PCIe H2D / D2H RDMA PCIe Root Complex P900 HBM NIC RDMA Host DRAM Host KV (HiCache) PCIe PCIe H2D / D2H RDMA NUMA interconnect remote NUMA access +1 hop, latency up, bandwidth down Access inside the same NUMA node (zero extra hops, ideal) Cross-NUMA remote access (+1 hop, latency up, bandwidth down)

Figure 1: P900 + HiCache server topology. Green = the P900 accesses host KV in its own NUMA node (the ideal path); red dashed = the P900 crosses the NUMA interconnect to reach another node's host KV (one extra hop, higher latency, lower effective bandwidth).

The division of labour, in one line each:

The core data path, one-line version:

P900 HBM -> PCIe -> PCIe Root -> NUMA Node -> Host DRAM -> Host KV

2. Five-Layer Optimization: From Topology to Actual Pinning

Five layers of NUMA optimization: keep the compute device close to the memory it reads 1 Identify the hardware topology numactl -H / lscpu -e / lspci -tv / cat numa_node, then draw the P900 to PCIe Root to NUMA table 2 Pin CPUs to the matching NUMA node rank0/1 to NUMA0, rank2/3 to NUMA2 (the idea behind flags like --numa-node 0 0 2 2) 3 Put host KV on the local NUMA node too CPU thread, host KV, PCIe Root and P900 should all agree on locality, not just the ranks 4 Consider NUMA for the NIC as well Under PD / Mooncake / RDMA, keep P900, host KV and NIC on the same NUMA node 5 Avoid over-pinning Balance locality against load; when requests are skewed, hard pinning can be worse Experiment matrix: correct NUMA > automatic NUMA > remote NUMA, showing HiCache performance is clearly sensitive to NUMA locality

Figure 2: Five layers of NUMA optimization. From "see the topology" through pinning CPUs, KV and NIC, to the final step of not over-pinning: find the balance between locality and load.

3. What You Should Measure Is Not “NUMA On / Off”

The valuable experiment compares different locality configurations, not a simple on/off switch:

ConfigurationPurpose
Correct topology bindingSee the upper bound of the locality win
Automatic NUMA (system assigned)Baseline for comparison
Wrong / remote bindingMeasure how large the NUMA penalty is
Shared KV + NUMA ONBest-case locality for shared KV
Shared KV + NUMA OFFHow sensitive shared KV is to NUMA

If the results come out as:

correct NUMA  >  automatic NUMA  >  remote NUMA

that is essentially proof that HiCache performance is clearly affected by NUMA locality — which is exactly why parameters like --numa-node 0 0 2 2 exist.

4. One-Line Summary

PCIe answers “how does the P900 exchange data with the host / CPU”; NUMA answers “which CPU memory node holds that host data and is it close enough to the P900”; and in PD / Mooncake setups the NIC also has to be nearby.

So what you are ultimately optimizing is an entire path:

P900  <->  PCIe Root Complex  <->  NUMA Node  <->  Host DRAM  <->  Host KV

not a single --numa-node flag.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。