NVIDIA Proposes AI Factory Architecture to Integrate Custom XPUs with Existing Infrastructure

NVIDIA has unveiled a blueprint for an AI Factory concept designed to help hyperscalers and AI-native enterprises integrate custom accelerators, or XPUs, into established AI infrastructures. Instead of treating hardware as a collection of individual chips, this design approach treats the AI infrastructure as a unified platform. It spans rack-scale architectures, scale-up and scale-out networking, and a complete software stack to optimize system-wide operations.
Related tools
Recommended tools for this topic
These picks prioritize high-intent tools relevant to this topic. Some links may include partner or affiliate tracking.
Strong fit for AI, backend, and frontend readers looking for an AI-first coding workflow.
View CursorStrong cloud alternative for startups and developer-led infrastructure decisions.
View DigitalOceanNatural next step for readers evaluating LLM adoption, APIs, and production inference.
Explore APIComparison
| Aspect | Before / Alternative | After / This |
|---|---|---|
| Design Philosophy | An aggregate of individual accelerators connected via generic interfaces | A unified rack-scale factory architecture designed for continuous operations |
| Optimization Metric | Peak theoretical FLOPS of individual computing chips | System-level output metrics including tokens per second per watt and cost per token |
| Custom Silicon Focus | In-house development of both custom compute cores and complex peripheral subsystems | Specialization in unique compute logic while relying on standardized NVLink Fusion integration |
| Thermal and Power Management | Standard generic server chassis limits and air cooling solutions | Co-designed rack-scale thermal management and advanced power distribution systems |
Action Checklist
- Evaluate the thermal and power delivery capacity of existing data centers against rack-scale integration standards AI Factory designs significantly exceed traditional commodity server power densities.
- Align custom XPU interconnect physical layers with NVLink Fusion specifications Ensures seamless scale-up and scale-out fabric connectivity across nodes.
- Analyze total cost of ownership using token-based efficiency metrics rather than simple chip-acquisition costs Measure performance in tokens per second per watt and system utilization rates.
- Assess software stack compatibility to ensure custom silicon integrates with existing deep learning frameworks and libraries A highly customized software layer is required to bridge proprietary XPUs with common open-source runtimes.
Source: NVIDIA
This page summarizes the original source. Check the source for full details.

