Mark RadarMARK RADAR
About
EN
Sign in

AMD Bets on Inference as CPU Demand Outpaces GPUs

2 reports · First detected 2026-07-24 · Last active 2026-07-24

AI computing has been dominated by the costly training of large models on GPUs, but deployment shifts the burden toward inference, where models respond continuously to users. Agentic AI adds code execution, API calls, data movement and task orchestration, expanding the role of CPUs alongside accelerators. That change gives Advanced Micro Devices a broader opening against Nvidia because AMD sells EPYC CPUs, Instinct GPUs, networking silicon and ROCm software as an integrated data-center stack.

At AMD’s Advancing AI 2026 event in San Francisco on July 23, CEO Lisa Su said inference would account for 60% of global AI compute in 2026 and CPU demand was growing even faster than GPU demand. AMD forecast the AI accelerator market at $1.4 trillion and the server CPU market above $200 billion by 2030. Its Helios rack links 72 Instinct MI455X accelerators with 31 terabytes of HBM4 memory, with shipments due by the end of the third quarter and a broader ramp in the fourth.

All Coverage

2 original reports

The Backstory

The history behind this event
After this
AMD, Cerebras Team Up to Accelerate AI Inferencefirst seen 2026-07-29 · 1 reports · similarity 0.76 · same topic: AMD

AI inference is the computing stage that turns a trained model into answers, making latency, throughput and power efficiency central to user experience and operating costs. AMD and Cerebras are pursuing a disaggregated architecture that assigns prompt prefill and long-context processing to AMD’s Helios rack-scale platform, while Cerebras’ Wafer-Scale Engine handles memory-intensive decoding and token generation. The approach gives AMD another route to challenge NVIDIA in data-center AI as demand shifts toward real-time agents and coding assistants.

AMD and Cerebras Systems announced the partnership on July 23, 2026, at Advancing AI 2026. The companies said the combined workflow could deliver as much as five times more tokens per second per watt, with Cerebras planning to deploy Helios systems in its data centers. The service is due to launch first through Cerebras Cloud in the second half of 2026. Separately, OpenAI researcher Jeffrey Wang said an internal model ran so quickly that he could barely switch screens, underscoring the productivity gains possible from faster inference.

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)