VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909726255611904 |
|---|---|
| author | Ma, Qianli Zheng, Yaowei Shi, Zhelun Zhao, Zhongkai Jia, Bin Huang, Ziyue Lin, Zhiqi Li, Youjie Yang, Jiacheng Peng, Yanghua Zhang, Zhi Liu, Xin |
| author_facet | Ma, Qianli Zheng, Yaowei Shi, Zhelun Zhao, Zhongkai Jia, Bin Huang, Ziyue Lin, Zhiqi Li, Youjie Yang, Jiacheng Peng, Yanghua Zhang, Zhi Liu, Xin |
| contents | Recent advances in large language models (LLMs) have driven impressive progress in omni-modal understanding and generation. However, training omni-modal LLMs remains a significant challenge due to the heterogeneous model architectures required to process diverse modalities, necessitating sophisticated system design for efficient large-scale training. Existing frameworks typically entangle model definition with parallel logic, incurring limited scalability and substantial engineering overhead for end-to-end omni-modal training. We present VeOmni, a modular and efficient training framework to accelerate the development of omni-modal LLMs. VeOmni introduces model-centric distributed recipes that decouples communication from computation, enabling efficient 3D parallelism on omni-modal LLMs. VeOmni also features a flexible configuration interface supporting seamless integration of new modalities with minimal code change. Using VeOmni, a omni-modal mixture-of-experts (MoE) model with 30B parameters can be trained with over 2,800 tokens/sec/GPU throughput and scale to 160K context lengths via 3D parallelism on 128 GPUs, showcasing its superior efficiency and scalability for training large omni-modal LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_02317 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo Ma, Qianli Zheng, Yaowei Shi, Zhelun Zhao, Zhongkai Jia, Bin Huang, Ziyue Lin, Zhiqi Li, Youjie Yang, Jiacheng Peng, Yanghua Zhang, Zhi Liu, Xin Computation and Language Artificial Intelligence Distributed, Parallel, and Cluster Computing Recent advances in large language models (LLMs) have driven impressive progress in omni-modal understanding and generation. However, training omni-modal LLMs remains a significant challenge due to the heterogeneous model architectures required to process diverse modalities, necessitating sophisticated system design for efficient large-scale training. Existing frameworks typically entangle model definition with parallel logic, incurring limited scalability and substantial engineering overhead for end-to-end omni-modal training. We present VeOmni, a modular and efficient training framework to accelerate the development of omni-modal LLMs. VeOmni introduces model-centric distributed recipes that decouples communication from computation, enabling efficient 3D parallelism on omni-modal LLMs. VeOmni also features a flexible configuration interface supporting seamless integration of new modalities with minimal code change. Using VeOmni, a omni-modal mixture-of-experts (MoE) model with 30B parameters can be trained with over 2,800 tokens/sec/GPU throughput and scale to 160K context lengths via 3D parallelism on 128 GPUs, showcasing its superior efficiency and scalability for training large omni-modal LLMs. |
| title | VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo |
| topic | Computation and Language Artificial Intelligence Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2508.02317 |