Bull Selected Over HPE for Next-Generation Lumi AI Supercomputer
Europe’s National Labs and AI Infrastructure Europe’s approach to generative AI infrastructure remains closely tied to its national supercomputing laboratories. These facilities...
By Software Development Team
Europe’s National Labs and AI Infrastructure
Europe’s approach to generative AI infrastructure remains closely tied to its national supercomputing laboratories. These facilities have built systems capable of both traditional high-performance computing (HPC) simulation and modeling, while efforts to customize generative AI and machine learning models are generally concentrated in the national labs rather than among AI model developers.
The largest AI model developers are based primarily in the United States and China. Across Europe, national laboratories are therefore leading both HPC and generative AI initiatives. They apply generative AI to government-funded scientific challenges while also providing compute allocations to European businesses developing AI models or using HPC tools to design products.
National laboratories are not financed by private equity firms, GPU accelerator vendors, or major cloud providers. Their budgets and computing resources are consequently much smaller, but they still have access to more compute capacity than in the past and support public-sector and commercial research programs.
The Lumi AI Factory
The next-generation Lumi-AI supercomputer will form part of what is being called the Lumi AI factory. The project is funded by the EuroHPC Joint Undertaking and will be installed at the Center for Scientific Computation (CSC) facilities in Kajaani, Finland. “Lumi” means “snow on the ground” in Finnish, although the system’s mascot is a white wolf.
The original Lumi hybrid supercomputer cluster was first unveiled in October 2020. It combines CPU-only, CPU-GPU, large-memory CPU, object storage, Lustre parallel file system storage, and flash storage nodes. The system was planned to deliver 552 petaflops of aggregate FP64 performance.
The Lumi-C partition, built with more than 200,000 AMD Epyc processor cores, was expected to provide about 2 petaflops of FP64 performance. The Lumi-G partition, based on AMD Instinct processors, was designed for approximately 550 petaflops of double-precision floating-point processing. The cluster also includes:
- 30 PB of Ceph object storage
- 80 PB of Lustre storage
- 7 PB of flash storage
- A 200 Gb/sec “Rosetta” Slingshot-10 interconnect
Hewlett Packard Enterprise was the main contractor for the original system. EuroHPC provided a budget of $237 million, equivalent to approximately $431 per teraflop. The Lumi-C partition was delivered in summer 2021, followed by the Lumi-G GPU partition at the end of 2021. Both partitions became operational in June 2022.
The Lumi-G partition ranked number 11 on the Top500 list in June, despite being four years old at the time.
Existing Lumi Partitions
Lumi-C contains 1,536 nodes. Each node has two 64-core “Milan” Epyc 7763 processors running at 2.45 GHz, for a total of 196,860 cores.
On the June Top500 list, Lumi-C ranked number 262. Its peak theoretical performance is 7.29 petaflops, while its sustained Linpack performance is 6.3 petaflops. That represents computational efficiency of 82.6 percent within a 1.22-megawatt power envelope, or 5.18 gigaflops per watt. The nodes are connected with a 200 Gb/sec Slingshot 11 interconnect.
Lumi-G contains 2,978 nodes. Each node uses one 64-core “Trento” Epyc 7A53 custom processor and four “Aldebaran” MI250X GPU accelerators. The Trento processor was developed for the “Frontier” supercomputer at Oak Ridge National Laboratory in the United States and provided memory coherence between AMD CPUs and GPUs before that capability became commercially available.
The partition contains:
- 2,978 Trento CPUs
- 190,592 CPU cores
- 11,912 Aldebaran GPUs
- 2.62 million compute units
- 167.7 million streaming processors
The nodes are connected by a 200 Gb/sec Slingshot-11 interconnect. Lumi-G has a theoretical peak of 531.5 petaflops at FP64 precision and achieved 379.7 petaflops on Linpack, for computational efficiency of 71.4 percent. It uses 7.1 megawatts and delivers 53.43 gigaflops per watt on Linpack, more than ten times the power efficiency of the CPU-only partition.
Expected Lumi-AI Performance
The Lumi-AI partition is expected to be an upgrade to Lumi-G. The existing Lumi-G system is expected to remain in production because of the value and limited availability of its AMD GPUs.
The new partition is expected to provide ten times the performance on AI workloads and approximately twice the performance on FP64 HPC workloads. Improving the AI performance of the MI250X is relatively straightforward because its INT4 and INT8 capabilities were implemented using FP16 resources. The performance was effectively the same across INT4, INT8, and FP16 formats.
The Altair GPUs used in the MI430X are designed for both HPC and AI workloads. Unlike the top-end Altair MI455X, which targets AI inference and training, the MI430X supports native FP4 and FP8 processing. Moving from FP16 to FP4 accounts for a fourfold performance effect within the projected tenfold AI improvement. The remaining 2.5-fold increase comes from the higher raw FP16 performance of the MI430X compared with the Aldebaran GPU in the MI250X.
The higher performance should allow Lumi-AI to use substantially fewer nodes, although each node is expected to consume more power and generate more heat.
Projected Node Configuration
The MI250X is rated at 47.9 teraflops on its vector units and 95.7 teraflops on its matrix units. The MI430X has a peak FP64 performance of 288 teraflops. AMD did not specify whether that figure referred to vector or matrix units, although the MI430X delivers 288 teraflops on both.
Using the vector FP64 figure for the Lumi-G calculation gives 570.6 petaflops: 47.9 teraflops multiplied by four GPUs per node and 2,978 nodes. This does not match the 531.5-petaflop Rpeak figure reported for the Linpack test.
A twofold increase in FP64 vector performance would require approximately 3,962 MI430X GPUs, or 991 nodes. That would reduce the GPU partition’s node count by 66.7 percent and produce a theoretical peak of 1.14 exaflops. Each node would consume approximately twice as much power, however: 5 kilowatts for the Venice-Altair 1x4 board compared with 2.5 kilowatts for the Trento-Aldebaran 1x4 board. At 991 nodes, the partition would consume approximately 5 megawatts, a 33 percent reduction compared with Lumi-G.
A separate calculation produces only an eightfold AI performance increase with that node count. Achieving a tenfold increase at four-bit precision would require 1,240 nodes with four MI430X GPUs each. This could indicate that CSC Finland will receive approximately 25 percent more FP64 vector performance than publicly stated.
The projected configuration is therefore 1,240 nodes containing 4,960 MI430X GPUs. That would represent a 58.4 percent reduction in GPU partition node count and system power consumption of approximately 6.2 megawatts. Bull, the primary contractor for Lumi-AI, and CSC Finland have both stated that AI performance will be ten times higher and HPC performance approximately twice as high.
Using a 256-core “Venice” Epyc 9006 processor, the Lumi-AI partition would contain 317,440 CPU cores. This would provide additional compute resources for coordinating the GPUs and increase the core count by 66.7 percent compared with the current Lumi-G partition.
Cost and Contractor Selection
The Lumi-AI partition has a price of €387.8 million. Converted to U.S. dollars, that is approximately $445.1 million. Against a projected peak FP64 performance of 1.43 exaflops, the cost is approximately $312 per teraflop at peak.
That is 27.6 percent lower than the $431-per-teraflop peak cost of the complete Lumi cluster when it was installed in 2021. It is not known whether the Lumi CPU-only partition and the various storage clusters will also be upgraded under the Lumi-AI contract. The comparison assumes that they will be, creating a system-level comparison with the original Lumi installation.
CSC Finland and EuroHPC selected Bull over HPE for the project. EuroHPC has emphasized European suppliers where possible. Slingshot is an Ethernet variant that could potentially be adapted to hardware compliant with the Ultra Ethernet Consortium. Bull’s BXI protocol is also being adapted for UEC hardware and its features. As a result, customers using BullSequana XH3500 systems may have a choice of switching technologies for scale-out clusters.
The Lumi-AI GPU partition is expected to be installed during the second half of 2027 in a new data center in Kajaani. Finland, Czechia, Denmark, Estonia, Norway, and Poland are contributing funding alongside the European Union’s EuroHPC Joint Undertaking. The specific contribution from each participant has not been disclosed.
Lumi-IQ Quantum Computing System
The Lumi AI factory will also include a quantum computing component called Lumi-IQ. The system is being built by IQM Quantum Computers, a Finnish company that has raised €600 million in venture capital and is seeking to become a leading European quantum computing supplier. IQM is preparing to go public and has a current valuation of $1.8 billion.
CSC Finland plans to install a Halocene H4 system in 2027. It will have 150 physical qubits and support five logical qubits. Upgrades planned for 2028 are intended to improve error correction and logical-qubit reliability.
In 2029, CSC Finland is expected to receive a Halocene H5 quantum computer. Its number of physical qubits has not been disclosed, but the system is expected to provide at least nine logical qubits. That could imply approximately 270 physical qubits for the logical-qubit capacity.