VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ma, Qianli, Zheng, Yaowei, Shi, Zhelun, Zhao, Zhongkai, Jia, Bin, Huang, Ziyue, Lin, Zhiqi, Li, Youjie, Yang, Jiacheng, Peng, Yanghua, Zhang, Zhi, Liu, Xin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909726255611904
author Ma, Qianli
Zheng, Yaowei
Shi, Zhelun
Zhao, Zhongkai
Jia, Bin
Huang, Ziyue
Lin, Zhiqi
Li, Youjie
Yang, Jiacheng
Peng, Yanghua
Zhang, Zhi
Liu, Xin
author_facet Ma, Qianli
Zheng, Yaowei
Shi, Zhelun
Zhao, Zhongkai
Jia, Bin
Huang, Ziyue
Lin, Zhiqi
Li, Youjie
Yang, Jiacheng
Peng, Yanghua
Zhang, Zhi
Liu, Xin
contents Recent advances in large language models (LLMs) have driven impressive progress in omni-modal understanding and generation. However, training omni-modal LLMs remains a significant challenge due to the heterogeneous model architectures required to process diverse modalities, necessitating sophisticated system design for efficient large-scale training. Existing frameworks typically entangle model definition with parallel logic, incurring limited scalability and substantial engineering overhead for end-to-end omni-modal training. We present VeOmni, a modular and efficient training framework to accelerate the development of omni-modal LLMs. VeOmni introduces model-centric distributed recipes that decouples communication from computation, enabling efficient 3D parallelism on omni-modal LLMs. VeOmni also features a flexible configuration interface supporting seamless integration of new modalities with minimal code change. Using VeOmni, a omni-modal mixture-of-experts (MoE) model with 30B parameters can be trained with over 2,800 tokens/sec/GPU throughput and scale to 160K context lengths via 3D parallelism on 128 GPUs, showcasing its superior efficiency and scalability for training large omni-modal LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02317
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
Ma, Qianli
Zheng, Yaowei
Shi, Zhelun
Zhao, Zhongkai
Jia, Bin
Huang, Ziyue
Lin, Zhiqi
Li, Youjie
Yang, Jiacheng
Peng, Yanghua
Zhang, Zhi
Liu, Xin
Computation and Language
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Recent advances in large language models (LLMs) have driven impressive progress in omni-modal understanding and generation. However, training omni-modal LLMs remains a significant challenge due to the heterogeneous model architectures required to process diverse modalities, necessitating sophisticated system design for efficient large-scale training. Existing frameworks typically entangle model definition with parallel logic, incurring limited scalability and substantial engineering overhead for end-to-end omni-modal training. We present VeOmni, a modular and efficient training framework to accelerate the development of omni-modal LLMs. VeOmni introduces model-centric distributed recipes that decouples communication from computation, enabling efficient 3D parallelism on omni-modal LLMs. VeOmni also features a flexible configuration interface supporting seamless integration of new modalities with minimal code change. Using VeOmni, a omni-modal mixture-of-experts (MoE) model with 30B parameters can be trained with over 2,800 tokens/sec/GPU throughput and scale to 160K context lengths via 3D parallelism on 128 GPUs, showcasing its superior efficiency and scalability for training large omni-modal LLMs.
title VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
topic Computation and Language
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2508.02317