AMD AI Efficiency: Rack Scale Key to 20×30 Goal
- Advanced Micro Devices (AMD) is pursuing an ambitious objective: a 20-fold increase in the energy efficiency of its chips by 2030.
- Sam Naffziger, AMD senior vice president and fellow, said that larger devices are more efficient.
- The MI300 series of APUs and GPUs exemplifies this strategy, using 3D stacking to create dense compute units.
AMD is aggressively targeting a 20x boost in chip energy efficiency by 2030, and rack-scale design forms the core of their strategy.This aspiring initiative, detailed in the article, tackles the dual challenges of soaring datacenter power demands and the plateauing of Moore’s Law. AMD is leveraging complex hardware-software co-design to maximize the performance gains. They are going beyond chiplets to integrate at the data center level. The MI400 rack-scale compute platform represents the next evolution, with photonic interconnects perhaps replacing copper. For further insights into this initiative, consider visiting News Directory 3. Discover what’s next in the quest for sustainable, high-performance computing.
AMD Aims for 20x Efficiency Boost with Rack-Scale Chip Design
Updated June 12, 2025
Advanced Micro Devices (AMD) is pursuing an ambitious objective: a 20-fold increase in the energy efficiency of its chips by 2030. This initiative addresses concerns about datacenter power consumption and the slowing pace of Moore’s Law. AMD sees rack-scale architectures as a vital component in achieving this goal, focusing on energy efficient computing.
Sam Naffziger, AMD senior vice president and fellow, said that larger devices are more efficient. He added that what used to require an entire rack of computing devices can now fit into a single package. this approach builds upon AMD’s earlier adoption of chiplet architectures in its CPUs and GPUs, which allowed the company to enhance performance per watt.
The MI300 series of APUs and GPUs exemplifies this strategy, using 3D stacking to create dense compute units. Now, AMD is extending its focus beyond individual chips to the rack level to further improve efficiencies.
Naffziger stated that architecting at the data center level is the way to deliver continued significant improvements. NVIDIA also revealed its rack-scale system, the GB200 NVL72, at GTC last year.
Historically, GPU systems from both companies have used high-speed interconnects to combine resources. NVIDIA’s GB200 NVL72 uses 18 NVLink switch chips to integrate 72 Blackwell GPUs into a single unit,consuming 120kW. NVIDIA plans to expand this architecture to support even more GPUs and power in the future.
Naffziger contends that rack scale is reinventing the scale-up multi-processing that IBM did in the 80s with shared memory spaces. Rather of a few dozen System/370 mainframes, the industry is now talking about tens, perhaps hundreds of GPUs.
AMD’s first rack-scale compute platform, the MI400, is expected next year. It will likely use the Global Accelerator link (UALink) interconnect,similar to NVIDIA’s NVLink. Future designs may incorporate photonic interconnects to replace copper, offering greater bandwidth, but also presenting challenges related to temperature sensitivity and mechanical robustness.
Naffziger said that economics drive everything, and the industry is at the point where economics will favor optical.
He also noted the temperature sensitivities with optical. He added that there’s a lot more to worry about than in electrical space and that the industry now has to route fiber attach and make sure it’s mechanically robust and not susceptible to vibration.
Process technology and advanced packaging will also contribute to AMD’s 20×30 goal.Naffziger said that there’s still the remnants of Moore’s law out there and that the company has to use the latest process nodes.
Improvements in memory technology, such as 3D stacking and HBM customization, can further reduce power consumption. Hardware-software co-design will also be crucial, as raw hardware gains are diminishing.
AMD has invested in optimizing its ROCm software stack and acquired companies like Nod.ai, Mipsology, and Brium to enhance its AI capabilities. Sharon Zhou, CEO of Lamini, recently announced her move to AMD to support these efforts.
Naffziger said that when the company talks about a rack-scale goal, there definitely are big opportunities in system architecture, system design, improved components, integration reducing the cost of communication, but that the company has to map the workload optimally on that hardware.
Support for lower precision datatypes like FP8 and FP4 is another example, trading output quality for a smaller memory footprint and increased throughput. However, software must adapt to these new data types.
The AI ecosystem’s rapid evolution poses challenges for performance measurement. Naffziger said that the company can’t assume Llama 405B is going to be here in 2030 and have any meaning.
AMD will use a combination of GPU FLOPS, HBM, and network bandwidth, weighted for inference and training, to track progress toward its 20×30 goal.
