d-Matrix announced a multi-year collaboration with NVIDIA on September 10 to connect its next-generation Raptor inference XPUs to NVIDIA’s AI infrastructure platform through NVLink Fusion. The announcement gives a specialized inference-chip company a stated path into NVIDIA’s rack-scale ecosystem, but it describes a roadmap rather than a shipping deployment.
A rack-scale path for custom inference silicon
NVIDIA says NVLink Fusion will connect custom XPUs and CPUs to its NVLink scale-up and Spectrum-X scale-out networking. d-Matrix’s release says Raptor will be designed into an NVIDIA MGX rack with Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet. The two official announcements match on the core integration, while d-Matrix’s release supplies the component-level roadmap.
The companies frame the architecture around latency-sensitive inference such as coding assistants, real-time chatbots, and voice agents. d-Matrix says operators could split inference stages between its XPUs and NVIDIA GPUs, but the cited material does not provide an independent test of latency, throughput, power, or cost.
The delivery date is still ahead
d-Matrix says Raptor is expected to tape out before the end of 2026 and that initial availability integrated into NVIDIA MGX racks is expected in the fourth quarter of 2027. Those are forward-looking company targets. The announcement does not name a customer deployment, disclose commercial pricing, or establish that the planned rack has entered production.
The concrete change is the announced integration path: custom inference silicon can be designed for a common NVIDIA rack, networking, power, cooling, and supply-chain framework. Whether that reduces deployment risk or improves economics for operators remains an open question until hardware and customer evidence are available.
