PCIe peer-to-peer transfer can reduce multi-GPU inference overhead by moving tensors directly between compatible GPU memories instead of staging them through host RAM.
A local model split across two GPUs must transfer activations, KV blocks, or expert outputs whenever execution crosses a device boundary. Without direct peer access, data may travel from one GPU to host memory and then back to the other, consuming CPU-facing links and adding copies. P2P shortens that path, but its value depends on topology, transfer size, synchronization, and how often the model communicates.
Peer Access Replaces a Host-Staged Copy Path
With peer access enabled, one GPU can address or copy data in another GPU's memory through the supported interconnect path. The transfer avoids an explicit bounce buffer in pageable or pinned system memory and can reduce CPU involvement.
The CUDA programming guide explains that peer memory access must be supported and enabled between device pairs. Capability is directional and topology-dependent, so software should query each pair rather than assuming every GPU in one host can communicate directly.
The model still needs synchronization so a consumer GPU does not read incomplete activations. P2P removes a staging step; it does not remove ordering, kernel launch, or collective-communication costs. This distinction remains visible during later household testing.
PCIe Topology Determines the Direct Path's Real Bandwidth
Two GPUs beneath the same PCIe switch can often exchange traffic without crossing a CPU socket, while devices behind different root complexes may require a host path or lose P2P support. Link generation, lane width, switch oversubscription, and simultaneous traffic set the ceiling.
NCCL documents that it prefers direct GPU communication when CUDA reports compatible GPUs, using PCIe or NVLink according to available topology. Its topology tools expose whether each device pair can use direct PCIe access. The intermediate result must remain inspectable before automation follows.
Small transfers can remain dominated by launch and synchronization latency, while large tensor transfers approach link bandwidth. A pipeline with frequent narrow boundaries may gain less than a design that communicates fewer, larger blocks. That boundary should be measured separately under realistic operating conditions.
Partitioning Determines Whether Faster Copies Matter
Tensor parallelism communicates within many layers, pipeline parallelism transfers activations between stage boundaries, and expert parallelism exchanges routed tokens. The same P2P link can therefore be lightly used or become the main limit depending on partition strategy.
NVIDIA's GPUDirect design discussion shows how PCIe topology placement and switch placement affect direct data movement relative to CPU-staged paths. The principle applies even though inference frameworks add their own collective and scheduling layers. The practical consequence appears when several sources compete for limited context.
The failure boundary is unsupported topology, IOMMU or virtualization restrictions, or communication that already exceeds the PCIe budget. Software then falls back to host staging or suffers link contention, and adding a second GPU can slow inference despite greater compute capacity.
Measure Every GPU Pair and Model Boundary
Map GPU, CPU socket, PCIe root, switch, link generation, width, P2P capability, and NUMA memory. Benchmark one-way and bidirectional peer copies plus host-staged copies across representative transfer sizes. This dependency should remain explicit in the final interface.
Connect the result to NUMA-aware placement. Profile compute time, communication time, synchronization, collective bandwidth, tokens per second, and p99 request latency for each model partition with P2P enabled and disabled. The result must therefore be checked against the original evidence.
Keep the multi-GPU plan only when end-to-end latency or capacity improves. If direct copies are fast but inference remains communication-bound, reduce partition crossings or choose a topology-aware placement rather than treating P2P support as sufficient.
Tech & AI HUB
More to Read

What Is Embedding Drift, and When Does a Private Search Index Need Rebuilding?
Decode model, preprocessing, corpus, and query drift; distinguish monitoring from incompatibility; and decide when a private index needs rebuilding.

What Is Tokenizer Compatibility, and Why Can It Break Model Switching?
Decode vocabulary identity, special-token semantics, chat templates, cached tokens, adapters, and compatibility checks for local model switching.

What Is Model Residency, and When Should a Local AI Service Keep Weights Loaded?
Decode weight residency, cache levels, cold starts, eviction, multiplexing, memory pressure, and when a home AI service should stay warm.

