Microsoft to Deploy AMD Helios Rack Systems on Azure for AI Inference

AMD’s first rack-scale AI platform, Helios, is heading to Azure. Microsoft announced Sunday that it will deploy the system across its cloud infrastructure, giving customers a new path for running AI inference on AMD silicon.

The move reflects how fast AI workloads are outpacing conventional infrastructure. No single approach can keep up anymore, Microsoft said, and the company has been aggressively upgrading Azure’s compute clusters to meet the demand. AMD’s latest chips are the newest addition to that expansion.

Three new Azure virtual machine families will run on the AMD hardware. HDv2 VMs target large-scale data processing. HXv2 VMs are built for electronic design automation — the chip design software that companies like AMD and Intel themselves rely on. The ND MI455X v7 VMs are designed specifically for AI inference, the phase where a trained model actually processes user requests in real time.

AMD first showed off Helios at Computex in Taipei back in June. It’s the company’s first integrated rack-level system, and the raw numbers are striking. Each rack can carry a 256-core EPYC Venice processor paired with 72 Instinct MI455X accelerators. That configuration delivers 31 terabytes of HBM4 memory and a theoretical 2,900 petaflops of FP4 dense compute.

AMD plans to start shipping Helios systems within the year. Microsoft, meanwhile, has been diversifying its AI infrastructure. It already developed a custom in-house AI chip, the Azure Maia 100, and earlier added AMD MI300X-powered VMs alongside its NVIDIA-based instances. The Helios deployment gives Azure customers another option — and reduces reliance on any single chip supplier.