NVIDIA Considers Reducing Memory in Rubin Ultra to Address High-Bandwidth Memory Supply Constraints

Deep News06:15

The artificial intelligence infrastructure boom is pushing supply chains to their limits.

NVIDIA is evaluating a significant adjustment to its next-generation AI GPU, the Rubin Ultra. The company is considering producing versions with lower memory configurations than originally planned, aiming to alleviate production pressure caused by a shortage of high-end High Bandwidth Memory (HBM).

This move indicates that even NVIDIA, which holds a dominant position in the GPU market, must balance product specifications with supply availability. The change reflects that HBM supply has become a critical bottleneck in the AI industry chain. It also suggests that AI server costs could rise further, and the pressure on data center construction will continue to be passed down the line.

On Thursday, NVIDIA shares closed down 0.1%, while SK Hynix shares fell 4.97%.

Addressing HBM Shortages with Multiple Lower-Memory Variants

Sources familiar with the matter said that over the past few weeks, NVIDIA has tested at least three different versions of the Rubin Ultra GPU. Some of these versions use less memory capacity than initially planned. A key reason for considering lower-memory versions is the potential inability to secure enough high-end HBM chips to support mass production of the original design.

The Rubin Ultra is a key component of NVIDIA's next-generation AI computing platform. It is positioned above the upcoming Rubin series and is seen as an important hardware platform for training ultra-large-scale AI models in the future. The original plan was to equip it with higher-capacity, higher-bandwidth HBM to improve model training and inference performance. However, with HBM supply remaining tight, NVIDIA is reassessing the product configuration. The goal is to ensure the product launches on schedule by adjusting memory specifications, rather than waiting for the supply chain to expand capacity.

Reduced Memory May Require Deploying More GPUs

For AI training, HBM not only determines a GPU's data throughput but also directly impacts the size of large models it can handle. If the Rubin Ultra ends up with a lower memory configuration, customers running large AI workloads, such as large language models, may need to deploy more GPUs to complete computing tasks that fewer chips could have handled originally.

While NVIDIA can partially compensate for the performance loss by increasing GPU compute power or optimizing interconnect bandwidth, overall system deployment costs and cluster complexity are likely to rise. For major cloud providers like Microsoft, Meta, Amazon, and Google, which are continuously increasing their AI capital expenditures, this implies further increases in future data center construction costs.

The AI Boom Exposes the Industry's Biggest Bottleneck

HBM has become one of the most critically scarce core components in the AI industry chain. In recent years, the rapid growth in demand for large model training has continuously pushed up GPU shipments. The HBM required for these GPUs is primarily supplied by a few companies, including Samsung Electronics, SK Hynix, and Micron. Due to the complex manufacturing process, high yield requirements, and long expansion cycles for HBM, supply growth has consistently struggled to keep pace with AI demand.

Industry experts widely expect HBM to remain in short supply for the next several years. This is a major reason why AI server prices remain high. NVIDIA's consideration of reducing the memory configuration for the Rubin Ultra reflects that even with the strongest bargaining power in the industry, its product planning is still constrained by HBM supply.

Cost Pressure Spreads Across the Tech Industry

The impact of rising HBM prices is not limited to AI servers. As key components like GPUs and HBM remain in short supply, hardware costs across the entire technology industry are being pushed up. Companies are increasing budgets for AI infrastructure, and more firms are being forced to raise capital expenditures to meet the demands of generative AI deployment.

At the same time, rising upstream chip costs are beginning to affect the consumer electronics sector. Some hardware manufacturers, including Apple, have already absorbed supply chain cost pressures by raising the prices of their end products. Market analysts believe that as long as AI computing demand continues to grow rapidly and the rate of new HBM capacity release remains limited, the competition for high-end memory resources will persist. HBM supply capacity will continue to be a key factor determining the pace of AI industry expansion.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment