Skip to main content
Back to Blog
AI/MLEnterpriseInnovation
4 September 20266 min readUpdated 10 September 2026

d-Matrix Plans to Pair Its Raptor Memory-Based XPU With Nvidia Rack-Scale Systems

d Matrix targets low latency AI inference Startup d Matrix has spent the past seven years assembling a hardware and software platform intended to accelerate AI inference while r...

By Hardware Team

d-Matrix targets low-latency AI inference

Startup d-Matrix has spent the past seven years assembling a hardware and software platform intended to accelerate AI inference while reducing operating costs. The company has focused on inference since its founding in 2019, when, according to co-founder and chief executive officer Sid Seth, relatively few people were familiar with the term.

Demand for faster inference has increased sharply over the past year. The growth of agentic AI tools such as Anthropic’s Claude Code, along with the introduction of OpenClaw, has added pressure to workloads that GPUs from Nvidia and AMD may not handle efficiently on their own.

“Over the last twelve months, there has been an explosion of low-latency inference across many different applications,” Seth said during a briefing. “Now, with the arrival of cybersecurity applications, low-latency tokens are in huge demand. Because of this explosion in low-latency inferencing, the demand for inference has really shot off the charts.”

After seven years of development and more than $500 million in funding, including investment from Microsoft’s M12 venture arm, d-Matrix introduced its first inference accelerator platform, Corsair, in June. Company executives said its memory-centric design could run inference workloads up to ten times faster than Nvidia GPUs alone.

The company is now pairing Corsair with Nvidia’s Blackwell GPU accelerators in a rack-scale system.

Memory-centric inference architecture

The XPU closely couples memory, including SRAM or 3D-RAM, with GPUs and CPUs in the same rack. The design is intended to accelerate token generation and reduce associated costs.

In a disaggregated inference environment, GPUs are suited to the compute-intensive prefill phase, during which context tokens are processed. Corsair, like Nvidia’s Groq 3 LPX accelerator, is intended for the decode phase, when an AI model generates code or other responses to queries from people or other models.

Following its April acquisition of GigaIO’s data center business, d-Matrix also introduced the SquadRack reference design. The design was developed with Arista Networks, Broadcom, and Supermicro.

“Our entire approach is predicated on doing more with the capital people deploy in our compute,” Seth said. “We are able to run really, really fast compute. We do more inference with a little amount of time, and we have made a very energy-efficient solution with memory-centric computing. That allows us to do more with less of these resources: money, time, and energy. We can hopefully, over time, alleviate the need to build out more data centers.”

Raptor and the 3D-RAM accelerator

Corsair has been available for only a few months, but d-Matrix is already discussing its follow-on platform, Raptor. The company detailed Raptor at Hot Chips 2026 and has placed Lightning, which will feature a multi-high DRAM stack, on its roadmap.

Raptor is scheduled to tape out by the end of 2026 and be released in the fourth quarter of 2027. Its 3D-RAM accelerator will use a 4-nanometer compute die manufactured by Taiwan Semiconductor Manufacturing Co. The compute die will be fused onto a custom DRAM die at a 36-micron pitch, delivering 100 TB/sec of bandwidth.

Raptor uses a variant of the chip-on-wafer package currently used for HBM packaging. The design combines a DRAM memory chip with an SRAM compute chip.

Integration with Nvidia MGX

A new partnership with Nvidia is intended to extend d-Matrix’s reach into enterprise and HPC data centers, AI laboratories, hyperscalers, large cloud providers, and neocloud companies. Announced Thursday, the partnership will make the Raptor XPU platform available through Nvidia’s MGX reference architecture as part of Nvidia’s AI factory offering.

Seth said Raptor needed a system in which it could be deployed. Although d-Matrix could have built its own liquid-cooled rack, Nvidia already provides an established ecosystem.

“This is truly the fastest way we feel of getting this amazing path-breaking technology that we are building at d-Matrix into the market with quick scale, with a robust supply chain to back it up and with a partner who is deployed in every data center across the world,” he said.

The partnership will also cover future d-Matrix XPU platforms, including Lightning.

Nvidia’s AI factory business is expanding. When Nvidia reported its Q2 2027 results, chief financial officer and executive vice president Colette Kress said the revenue opportunity for the AI factory platform had grown from approximately $18 billion per gigawatt with the Grace-Hopper rack-scale systems in 2022 to $40 billion per gigawatt with the forthcoming Vera-Rubin systems.

Raptor in the NVL144 MGX rack

NVLink and NVSwitch are central to d-Matrix’s plans. Nvidia’s high-speed interconnect technology links GPUs, CPUs, and other accelerators within its rack-scale MGX architecture. The platform includes Vera CPUs, Rubin GPUs, BlueField-4 DPUs, Spectrum-X Ethernet networking, and ConnectX-9 SuperNICs.

The d-Matrix configuration will use Nvidia’s NVLink compute tray, but instead of Vera-Rubin chips, the tray will contain Raptor XPUs along with Nvidia Vera CPUs and BlueField DPUs. ConnectX and Spectrum-X networking will provide scale-out connectivity.

The trays will plug into the liquid-cooled NVL144 MGX rack, which will also use an NVLink switch tray. Each rack will contain 144 Raptor XPUs.

Each card will contain 2.3 TB of 3D-stacked DRAM operating at 100 TB/sec. Across a single rack, the system will provide aggregate memory bandwidth of 7.2 petabytes per second. The Raptor rack can also operate as a companion system to Nvidia racks running Vera-Rubin GPUs.

“The beauty of this solution is we take the Raptor trays, plug them into the same NVL144 MGX rack architecture, which is widely deployed across many data centers, and we get instant access to those data centers,” Seth said.

Reported model performance

At Hot Chips, d-Matrix presented results from two AI models running on a rack of Raptor chips: Z.ai’s GLM 5.2 and Moonshot AI’s Kimi K3.

  • GLM 5.2 achieved approximately 3,000 tokens per second per user.
  • Kimi K3 reached 1,000 tokens per second per user.

Seth said the vendor can scale the system across eight racks.

Jesse Clayton, principal product marketing manager for Nvidia’s Data Center GPU business, said the partnership reflects Nvidia’s architectural approach. He described the platform as “completely fungible” and “vertically integrated, but horizontally open.”

Under this approach, vendors can use selected parts of Nvidia’s technology stack, which also includes the Groq 3 LPX accelerator rack, while integrating their own products. NVLink Fusion ports allow other silicon to connect to the platform.