Skip to main content
Back to Blog
AI/MLEnterpriseProduct Development
6 October 20268 min readUpdated 8 October 2026

Memory Has Become the IT Industry’s Key Constraint, Especially for AI

For most of the history of high end commercial and technical computing, system attention focused on compute. CPUs dominated for roughly five decades, followed by GPUs during the...

By Software Development Team

For most of the history of high-end commercial and technical computing, system attention focused on compute. CPUs dominated for roughly five decades, followed by GPUs during the past decade and a half. Memory, storage, and I/O remained important, but they were usually treated as secondary concerns.

System and cluster budgets therefore emphasized maximizing compute capacity. Organizations sometimes underinvested in memory and I/O, even when doing so reduced CPU or GPU utilization.

The rise of machine learning, followed by generative AI and agentic AI, has changed that balance. Memory has moved to the center of system design and, across much of the IT sector, has become a more important constraint than compute. The scale and speed of this change have been particularly significant.

Semiconductor revenue shifts toward memory

Gartner’s latest chip revenue breakdown shows memory sales, including flash, compared with non-memory semiconductor sales:

The overall semiconductor market appears to be following a kind of Moore’s Law in revenue, with sales nearly doubling between 2025 and 2026. Growth is forecast to slow between 2026 and 2027. Competition eventually increases, while organizations also reach a point where they have enough capacity for their workloads.

The 2026 growth rates published by Gartner for DRAM and NAND flash can be used to estimate the split between the two categories. Other memory types are assumed to generate approximately $4 billion annually and to grow at twice the rate of global gross domestic product, a pattern often associated with legacy components and platforms.

Micron’s latest total addressable market forecast for HBM can also be used to estimate other DRAM categories, including DDR4, DDR5, LPDDR4, and LPDDR5. The resulting estimate indicates that these other DRAM products will continue to generate far more revenue than HBM, probably for the foreseeable future.

Demand and pricing across memory markets

Demand for DRAM main memory, HBM stacked memory, and flash storage has risen sharply. Commercial companies have increasingly adopted flash instead of disk drives to improve workload performance, while generative AI has created additional demand for both memory and storage.

The industry is also in a broad system upgrade cycle. Server fleets contain machines that are five, six, or seven years old and require replacement or expansion. AI systems are being designed with high core density to conserve space and power, which increases demand for high-capacity DRAM and flash.

Traditional IT workloads, including transaction processing, data warehousing, data analytics, and web infrastructure, now compete with large generative AI deployments for DRAM and flash supplies. DRAM availability is tightening as manufacturers redirect capacity toward HBM stacked memory for AI accelerators. The shift is intended to pursue higher revenue, although it does not necessarily produce higher profits.

These conditions have driven DRAM prices higher. In November 2022, when OpenAI’s ChatGPT helped initiate the generative AI boom, server DDR4 memory cost approximately $3 per GB, while DDR5 cost about $4 per GB. These figures represented average prices for midrange capacities. Denser memory cost substantially more per GB, while lower-capacity products cost less.

A 30 TB TLC enterprise SSD cost approximately 10 cents to 13 cents per GB at that time. Four years later, street prices for DDR4 server memory are approximately two to three times higher. DDR5 server memory is approximately nine to thirteen times more expensive, and a 30 TB TLC SSD costs about six to seven times more.

Part of the change reflects the comparison with a memory and flash price crash that began in late 2021. Oversupply has affected the memory and flash industries repeatedly over several decades.

Disk drive prices have also increased. A 30 TB nearline disk drive costs approximately 2.5 times as much as it did over the same period. Low-cost disk capacity is also difficult to secure because hyperscalers and cloud builders have contracted most of the supply from Seagate, Western Digital, and Toshiba, the three remaining disk drive manufacturers.

HBM pricing and capacity decisions

HBM differs from standard DRAM in several respects. Samsung, SK Hynix, and Micron Technology are the three HBM stacked-memory manufacturers. There are no open spot prices because the market is dedicated to GPU makers Nvidia and AMD, as well as XPU makers and their chip partners. As a result, pricing information is limited.

When HBM2 debuted in 2017, its reported price was approximately $20 per GB. That figure declined to about $8 per GB two years later. HBM2E, which increased bandwidth, cost approximately $10 per GB. HBM3 reached around $12 per GB, while HBM3E cost approximately $13.50 per GB.

HBM4 is expected to begin shipping with the next generation of GPUs and XPUs. Its reported price is approximately $16 per GB, with Nvidia’s contract price rumored to be between $560 and $600 for a 36 GB stack. Some expectations put HBM4E, with its substantial bandwidth increase, at twice the price per GB of HBM4. The progression from HBM2 through HBM4 suggests that approximately $20 per GB may be plausible, while $36 per GB would represent a much larger increase.

The finished HBM component price has increased by approximately 1.6 times since the generative AI boom began. At the same time, each manufacturing generation produces more discarded DRAM chips because normal yields remain below 100 percent. HBM manufacturers may raise HBM4E prices in 2027 to bring its profitability closer to that of standard DRAM.

The HBM costs cited above are manufacturing or contract prices, not the prices paid by end users for GPUs and XPUs. The HBM component represents approximately half of the end-user cost of an Nvidia or AMD GPU, and probably about 65 percent of the cost of an XPU accelerator produced by hyperscalers, cloud builders, or AI model developers. OpenAI has also entered this area with its “Jalapeno” AI accelerator.

GPU and XPU designers are evaluating ways to deliver more capacity and bandwidth from the HBM they can obtain. Some may continue using additional HBM3E stacks instead of moving to HBM4. This preserves much of the bandwidth, which is important during inference decoding, while reducing memory capacity.

Nvidia is rumored to have revised its original plans for the “Rubin Ultra” GPU, expected in 2027. The initial design was reported to include four reticle-limited GPU chiplets and up to 1 TB of HBM4E memory. AMD is maintaining 488 GB for its “Altair” MI455X GPU accelerator, while Nvidia is reportedly considering as little as 192 GB for some future Rubin and Rubin Ultra designs.

Custom HBM base dies

One likely development is the movement of the HBM controller from the GPU chiplet to a customizable base die. HBM4 and HBM4E make this possible. Nvidia calls its implementation NVHBM, although custom base dies are not exclusive to Nvidia.

Other vendors are expected to develop their own base dies. Some may also place mathematical functions in the base die, following the processor-in-memory and processing-in-memory approaches explored for approximately a decade, though those approaches have not gained significant market traction. Such functions could perform selected operations in memory instead of moving data across the memory channel to the GPU or XPU.

Nvidia states that integrating the memory controller into the 3D HBM stack rather than the XPU can provide up to 30 percent greater memory bandwidth, 15 percent lower HBM power consumption, and up to 25 percent more area on the XPU compute die compared with standard HBM4E.

That additional chip area could be used to increase performance across two chiplets. One possible configuration for Rubin Ultra would provide 512 GB of HBM4E, or perhaps 384 GB with the same bandwidth. The original Rubin Ultra design reportedly used four stacks per GPU chiplet. Each stack was sixteen chips high, with 4 GB per chip, producing 1,024 GB of capacity and, based on Samsung’s stack specifications, approximately 58 TB per second of aggregate memory bandwidth.

Reducing stack heights and cutting the chiplet count in half could maintain 28 TB per second of bandwidth while providing 256 GB with eight-high stacks, 384 GB with twelve-high stacks, or 512 GB with sixteen-high stacks.

GPU and XPU manufacturers are facing similar trade-offs. Their objective is to maximize HBM capacity and bandwidth for leading-edge devices while reducing the cost per gigabyte.