AMD has pulled back the curtain on Helios, its first AI system designed around an entire rack: 72 Instinct MI455X accelerators, 31 terabytes of HBM4 memory, and 2.9 exaflops of FP4 compute.
Eighteen compute trays, each holding four GPUs based on the new CDNA 5 architecture, form a building block with a clear goal: taking on Nvidia’s Vera Rubin NVL72. This is not a card but a pre-assembled compute unit that marks the culmination of a clear trajectory — the industry is redefining the elementary “node” for training and inference as a densely interconnected rack, rather than a single server with a handful of GPUs.
The stakes go beyond a technological chase between two manufacturers. Helios says something deeper to anyone evaluating on-premise deployments of large language models. Nvidia has dominated this segment with platforms like HGX and DGX, building an ecosystem that practically turned GPU demand into a near-obligatory choice. AMD’s entry, with an integrated offering and a memory pool that closes the gap with the green alternatives, begins to erode that positional rent. For research labs, regulated companies, and service providers that want to keep both data and hardware in-house, having at least two credible suppliers at rack scale changes bargaining power and, more importantly, infrastructure planning.
The abundance of HBM4 — roughly over 430 gigabytes of high-bandwidth memory per GPU, based on the total figure — is the real pivot. With workloads like inference on hundred-billion-parameter models, the ability to keep weights and caches in fast memory without resorting to complex distribution strategies across separate nodes is what separates acceptable latency from permanent bottlenecks. Such density, packed into a standard rack, also lowers cooling and power complexity compared to do-it-yourself setups with multiple disconnected chassis. For those already experienced in self-hosting, the message is clear: the rack is becoming the atom of compute, not the server.
Then there is a strategic signal about the role of software. AMD has historically struggled against CUDA, yet the emphasis on open ecosystems like ROCm and the growing support from leading inference frameworks — from vLLM to TensorRT-LLM with alternative backends — are reducing friction. The fact that Helios arrives with CDNA 5, purpose-built to natively accelerate transformer workloads and low-precision operations, suggests that AMD doesn’t just want to sell raw silicon but to deliver a system ready for production inference, where throughput per watt and memory density are the decisive metrics.
It’s not only hyperscalers who stand to benefit. Even smaller organizations that have so far given up on running large models on their own due to a lack of Nvidia alternatives might revisit their TCO calculations. Not to mention that digital sovereignty pushes governments and critical sectors to seek diverse hardware stacks, perhaps combining different vendors in dedicated clusters. Helios doesn’t rewrite the rules overnight, but it accelerates the notion that rack-level competition is now a fact, not a distant possibility.
💬 Comments (0)
🔒 Log in or register to comment on articles.
No comments yet. Be the first to comment!