NVIDIA Enables Integration of Custom XPUs into NVLink Infrastructure for AI Factories

NVIDIA has unveiled its latest approach to integrating proprietary custom XPUs, developed by hyperscalers and AI-native enterprises, into its established NVLink infrastructure. This integration shifts the focus from individual chip designs to managing scale-up networks, rack-scale architectures, and production-grade software under a unified AI Factory model. Unlike traditional setups where custom accelerators are deployed in isolated silos, the unified model fuses computing resources using NVLink to directly improve operational metrics such as tokens generated per second and energy efficiency per watt. Organizations with custom silicon can leverage this validated ecosystem to deploy large-scale training and inference environments while minimizing time-to-market complexity and engineering overhead. However, maximizing the performance of this unified environment requires comprehensive optimization across the entire platform, including rack designs and supplier ecosystems. Engineering teams must evaluate how their custom silicon specifications align with NVIDIA's full-stack software and networking requirements to resolve potential integration boundaries before deployment.
Related tools
Recommended tools for this topic
These picks prioritize high-intent tools relevant to this topic. Some links may include partner or affiliate tracking.
Strong cloud alternative for startups and developer-led infrastructure decisions.
View DigitalOceanHigh-value hosting and deployment path for frontend and cloud readers.
View VercelA strong security and edge platform match across CDN, Zero Trust, and app protection.
View CloudflareComparison
| Aspect | Before / Alternative | After / This |
|---|---|---|
| Integration scope | Isolated custom silicon deployments focusing on chip-level metrics | Rack-scale system integration combining networking, power, and cooling |
| Interconnect fabric | Proprietary or fragmented custom interconnects | Unified NVLink scale-up and scale-out networking |
| Primary optimization metric | Peak theoretical FLOPS per individual accelerator | Tokens generated per second and token efficiency per watt |
| Software environment | In-house custom software stacks requiring manual maintenance | NVIDIA-validated full-stack production software ecosystem |
Source: NVIDIA
This page summarizes the original source. Check the source for full details.

