NC / HOME NEWS

d-Matrix Links Raptor XPUs to NVIDIA NVLink Fusion

d-Matrix says its next-generation Raptor inference XPUs will use NVIDIA NVLink Fusion and MGX racks to reach larger AI deployments.

Official technical rendering of d-Matrix Raptor compute trays and an MGX rack with NVIDIA NVLink Fusion
Image: NVIDIA, d-Matrix NVLink Fusion announcement

d-Matrix says its next-generation Raptor inference XPUs will connect to NVIDIA’s rack-scale AI infrastructure through NVLink Fusion. NVIDIA reported the partnership in its September 10 announcement, while d-Matrix published its own announcement about the rack-scale deployment. The plan gives d-Matrix a route to combine custom inference silicon with NVIDIA networking, rack designs, and systems instead of building a separate platform around every processor family.

NVIDIA describes NVLink Fusion as the connection between custom XPUs or CPUs and the wider NVIDIA AI platform. For d-Matrix, that means Raptor can be designed into an MGX rack architecture and connected through NVIDIA’s scale-up and scale-out systems. The companies are talking about future deployment, not announcing that Raptor racks are already shipping to customers.

From a processor to a complete rack

The attraction is the infrastructure around the chip. NVIDIA says d-Matrix plans to use its MGX rack architecture, NVLink scale-up networking, and Spectrum-X Ethernet scale-out networking with the Raptor platform. The company also plans to integrate NVIDIA Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs, and Spectrum-X networking. Those components cover host processing, data movement, storage paths, and the fabric that ties the rack together.

That broader stack matters because an inference accelerator is not deployed in isolation. A data center operator still needs power delivery, cooling, rack certification, software support, supply-chain capacity, and a way to move requests and model data between processors. NVIDIA’s pitch is that a common rack platform can support GPUs, CPUs, and XPUs without requiring a completely separate architecture for each one.

The idea builds on NVIDIA’s effort to make its AI factories less dependent on a single processor type. Neon Control previously covered the company’s NVLink Fusion relationship with MediaTek, which showed the same platform opening to another custom silicon effort. The d-Matrix announcement is narrower and more specific: it centers on Raptor’s role in low-latency inference and its planned fit inside NVIDIA’s rack ecosystem.

Inference specialization inside a common platform

d-Matrix co-founder and CEO Sid Sheth said the company is trying to answer rising inference demand while limiting the capital and energy burden of deployment. In practical terms, the partnership lets d-Matrix concentrate on its XPU architecture while using validated NVIDIA systems and interconnects around it. NVIDIA says Raptor racks will also be able to work alongside GPU-based Vera Rubin NVL72 systems for disaggregated inference.

That combination could let operators match different processors to different stages of an AI service, but the announcement does not provide a performance result, customer shipment date, price, or production volume for Raptor. Those details remain open. The confirmed development is the planned integration itself and the list of NVIDIA infrastructure components d-Matrix intends to use.

For readers tracking the hardware underneath AI services, the move is another example of the industry separating compute engines from the racks and networks that make them useful. Our memory-bandwidth explainer looks at one part of that system constraint. d-Matrix’s next milestone will be turning this platform agreement into a working, supportable rack that customers can actually deploy.