AMD: 428-billion-parameter MiniMax-M3 model serves faster than NVIDIA B200 and B300 on Instinct MI355X GPUs
AMD has demonstrated serving the 428-billion-parameter MiniMax-M3 model on Instinct MI355X GPUs using the ATOM and ATOMesh software. Sparse attention cuts compute per token to one-twentieth of its predecessor, and results outperform NVIDIA B200 on a single node and NVIDIA B300 in a multi-node configuration.
This article was generated using artificial intelligence from primary sources.
What is MiniMax-M3 and how does AMD serve it?
MiniMax-M3 is a large MoE (mixture-of-experts) model with 428 billion parameters, of which only 22 billion are active per individual token, with native support for text, image, and video in the same model. AMD has demonstrated serving this model on Instinct MI355X GPUs using ATOM, a single-node engine for serving on a single node, and ATOMesh, a layer for orchestrating serving across multiple nodes at once.
How much does sparse attention reduce compute?
MiniMax Sparse Attention, a technique that lets the model process only the relevant parts of the input context instead of all tokens equally, reduces the compute required per token to one-twentieth (1/20) of the previous generation at a context length of one million tokens.
Results outperform NVIDIA B200 and B300
On a single node (ATOM EAGLE3 FP4), the system operates at two working points. At high interactivity of 340 to 370 tokens per second per user, it reaches about 0.6 thousand tokens/s per GPU, compared to roughly 0.2 thousand for NVIDIA B200. At high throughput (17 to 135 tokens/s per user), it reaches 3,600 to 8,500 tokens/s per GPU, while NVIDIA B200 achieves 2,500 to 7,700. In a multi-node configuration, ATOMesh paired with Mooncake outperforms even NVIDIA B300, AMD’s directly competing GPU segment.
Frequently Asked Questions
- What is MiniMax-M3?
- MiniMax-M3 is a large MoE (mixture-of-experts) model with 428 billion parameters, of which 22 billion are active per token, with native support for text, image, and video.
- What are ATOM and ATOMesh?
- ATOM is AMD's single-node engine for serving AI models on a single node with GPUs, while ATOMesh is a layer for orchestrating model serving across multiple nodes at once.
- How much faster is MiniMax-M3 with sparse attention?
- MiniMax Sparse Attention reduces the compute required per token to one-twentieth (1/20) of the previous generation at a context length of one million tokens, and serving on AMD Instinct MI355X GPUs outperforms the results of NVIDIA B200 and B300 GPUs.
Sources
📬 AI news in your inbox
A daily digest built your way — pick topics, sources and cadence. One-click unsubscribe.
Related news
NVIDIA: Spectrum-6 Ethernet switch delivers 102.4 Tb/s and double the throughput for gigascale AI factories
AMD: GEAK v3 Open-Source Agent Speeds Up GPU Kernels 3.02× on Instinct MI300X, MI355X, and RDNA4
NVIDIA: Vera Rubin Platform Trains the Largest AI Models with a Quarter of Blackwell-Generation GPUs Thanks to Codesign