AMD Partners With Cerebras to Challenge Nvidia in AI Inference
- AMD is partnering with chip startup Cerebras to implement "disaggregated inference," a strategy that splits AI workloads across different hardware types to improve efficiency.
- Under the agreement, AMD will integrate its Helios server system into Cerebras data centers later this year.
- The Helios system is designed to process high volumes of requests, while Cerebras utilizes a wafer-sized chip specialized in generating near-instantaneous responses.
AMD is partnering with chip startup Cerebras to implement “disaggregated inference,” a strategy that splits AI workloads across different hardware types to improve efficiency. Announced Thursday by CEO Lisa Su, the move aims to challenge Nvidia’s dominance by separating the process of handling AI prompts from the generation of responses.
Under the agreement, AMD will integrate its Helios server system into Cerebras data centers later this year. This architectural shift moves away from traditional setups where a single piece of hardware manages both the initial processing of a request and the final output generation, which AMD argues are fundamentally different tasks.
The Helios system is designed to process high volumes of requests, while Cerebras utilizes a wafer-sized chip specialized in generating near-instantaneous responses. This division of labor is the core of the disaggregated inference approach.
AMD Helios performance claims against Nvidia
At the Advancing AI event, AMD positioned the Helios server system as a direct competitor to Nvidia’s Vera Rubin NVL72 rack. The company claims that Helios provides up to 30% more inference tokens per dollar than the Nvidia equivalent.
Lisa Su stated during the event that Every Helios can deliver more performance for the largest models, more capacity for longer context, and the bandwidth to scale across thousands of racks
.
This push for market share comes as AI companies transition their focus from training large models to deploying them for active use, increasing the demand for inference-optimized hardware from providers like AMD, Nvidia, and Broadcom.
Industry shift toward disaggregated architectures
The partnership with Cerebras aligns with a broader trend in the semiconductor industry. According to a June report from UBS, the limitations of current hardware architectures are driving a shift toward disaggregated inference.
UBS noted that other major players are pursuing similar strategies to reduce costs and increase efficiency, specifically naming Amazon Web Services and Nvidia, the latter of which integrated AI hardware startup Groq.
However, UBS identified a primary technical hurdle in this transition: orchestration. The report stated that getting different types of chips to work together seamlessly presents new challenges for developers and operators.
AMD infrastructure partnerships and client base
AMD is expanding its footprint among the largest AI labs and cloud providers. The company announced a multibillion-dollar infrastructure partnership with Anthropic on Wednesday.
- Microsoft
- Meta
- OpenAI
- Oracle
By bundling various types of AI chips within the Helios system, AMD is attempting to provide a scalable alternative to the integrated racks currently dominated by Nvidia in the AI training and inference market.
