Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917123662544896 |
|---|---|
| author | Anthony, Quentin Tokpanov, Yury Szot, Skyler Rajagopal, Srivatsan Medepalli, Praneeth Golubeva, Anna Shyam, Vasu Washbourne, Robert Iyer, Rishi Chaurasia, Ansh Figliolia, Tomas Yang, Xiao Sarje, Abhinav Thorstensen, Drew Pearson, Amartey Grossbart, Zack van Patten, Jason Barsoum, Emad Gu, Zhenyu Fu, Yao Millidge, Beren |
| author_facet | Anthony, Quentin Tokpanov, Yury Szot, Skyler Rajagopal, Srivatsan Medepalli, Praneeth Golubeva, Anna Shyam, Vasu Washbourne, Robert Iyer, Rishi Chaurasia, Ansh Figliolia, Tomas Yang, Xiao Sarje, Abhinav Thorstensen, Drew Pearson, Amartey Grossbart, Zack van Patten, Jason Barsoum, Emad Gu, Zhenyu Fu, Yao Millidge, Beren |
| contents | We report on the first large-scale mixture-of-experts (MoE) pretraining study on pure AMD hardware, utilizing both MI300X GPUs and Pollara networking. We distill practical guidance for both systems and model design. On the systems side, we deliver a comprehensive cluster and networking characterization: microbenchmarks for all core collectives (all-reduce, reduce-scatter, all-gather, broadcast) across message sizes and GPU counts over Pollara. To our knowledge, this is the first at this scale. We further provide MI300X microbenchmarks on kernel sizing and memory bandwidth to inform model design. On the modeling side, we introduce and apply MI300X-aware transformer sizing rules for attention and MLP blocks and justify MoE widths that jointly optimize training throughput and inference latency. We describe our training stack in depth, including often-ignored utilities such as fault-tolerance and checkpoint-reshaping, as well as detailed information on our training recipe. We also provide a preview of our model architecture and base model - ZAYA1 (760M active, 8.3B total parameters MoE, available at https://huggingface.co/Zyphra/ZAYA1-base) - which will be further improved upon in forthcoming papers. ZAYA1-base achieves performance comparable to leading base models such as Qwen3-4B and Gemma3-12B at its scale and larger, and outperforms models including Llama-3-8B and OLMoE across reasoning, mathematics, and coding benchmarks. Together, these results demonstrate that the AMD hardware, network, and software stack are mature and optimized enough for competitive large-scale pretraining. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_17127 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design Anthony, Quentin Tokpanov, Yury Szot, Skyler Rajagopal, Srivatsan Medepalli, Praneeth Golubeva, Anna Shyam, Vasu Washbourne, Robert Iyer, Rishi Chaurasia, Ansh Figliolia, Tomas Yang, Xiao Sarje, Abhinav Thorstensen, Drew Pearson, Amartey Grossbart, Zack van Patten, Jason Barsoum, Emad Gu, Zhenyu Fu, Yao Millidge, Beren Computation and Language Artificial Intelligence Distributed, Parallel, and Cluster Computing We report on the first large-scale mixture-of-experts (MoE) pretraining study on pure AMD hardware, utilizing both MI300X GPUs and Pollara networking. We distill practical guidance for both systems and model design. On the systems side, we deliver a comprehensive cluster and networking characterization: microbenchmarks for all core collectives (all-reduce, reduce-scatter, all-gather, broadcast) across message sizes and GPU counts over Pollara. To our knowledge, this is the first at this scale. We further provide MI300X microbenchmarks on kernel sizing and memory bandwidth to inform model design. On the modeling side, we introduce and apply MI300X-aware transformer sizing rules for attention and MLP blocks and justify MoE widths that jointly optimize training throughput and inference latency. We describe our training stack in depth, including often-ignored utilities such as fault-tolerance and checkpoint-reshaping, as well as detailed information on our training recipe. We also provide a preview of our model architecture and base model - ZAYA1 (760M active, 8.3B total parameters MoE, available at https://huggingface.co/Zyphra/ZAYA1-base) - which will be further improved upon in forthcoming papers. ZAYA1-base achieves performance comparable to leading base models such as Qwen3-4B and Gemma3-12B at its scale and larger, and outperforms models including Llama-3-8B and OLMoE across reasoning, mathematics, and coding benchmarks. Together, these results demonstrate that the AMD hardware, network, and software stack are mature and optimized enough for competitive large-scale pretraining. |
| title | Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design |
| topic | Computation and Language Artificial Intelligence Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2511.17127 |