InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Qiaoling, Gu, Diandian, Wang, Guoteng, Chen, Xun, Xiong, YingTong, Huang, Ting, Hu, Qinghao, Jin, Xin, Wen, Yonggang, Zhang, Tianwei, Sun, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
by: Gu, Diandian, et al.
Published: (2024)
by: Gu, Diandian, et al.
Published: (2024)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
by: Chen, Qiaoling, et al.
Published: (2023)
by: Chen, Qiaoling, et al.
Published: (2023)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
by: Chen, Qiaoling, et al.
Published: (2025)
by: Chen, Qiaoling, et al.
Published: (2025)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
by: Chen, Qiaoling, et al.
Published: (2026)
by: Chen, Qiaoling, et al.
Published: (2026)
Characterization of Large Language Model Development in the Datacenter
by: Hu, Qinghao, et al.
Published: (2024)
by: Hu, Qinghao, et al.
Published: (2024)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
by: Chen, Qiaoling, et al.
Published: (2025)
by: Chen, Qiaoling, et al.
Published: (2025)
DynaShard: Secure and Adaptive Blockchain Sharding Protocol with Hybrid Consensus and Dynamic Shard Management
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
by: Qing, Yuhao, et al.
Published: (2025)
by: Qing, Yuhao, et al.
Published: (2025)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile
by: Zhang, Ruisi, et al.
Published: (2024)
by: Zhang, Ruisi, et al.
Published: (2024)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
by: Chen, Chang, et al.
Published: (2025)
by: Chen, Chang, et al.
Published: (2025)
SP-Chain: Boosting Intra-Shard and Cross-Shard Security and Performance in Blockchain Sharding
by: Li, Mingzhe, et al.
Published: (2024)
by: Li, Mingzhe, et al.
Published: (2024)
TorchGT: A Holistic System for Large-scale Graph Transformer Training
by: Zhang, Meng, et al.
Published: (2024)
by: Zhang, Meng, et al.
Published: (2024)
TRAIL: Cross-Shard Validation for Cryptocurrency Byzantine Shard Protection
by: Jacovetty, Mitch, et al.
Published: (2024)
by: Jacovetty, Mitch, et al.
Published: (2024)
ShardTensor: Domain Parallelism for Scientific Machine Learning
by: Adams, Corey, et al.
Published: (2026)
by: Adams, Corey, et al.
Published: (2026)
Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
by: Wang, Zerui, et al.
Published: (2025)
by: Wang, Zerui, et al.
Published: (2025)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
by: Jia, Jinda, et al.
Published: (2024)
by: Jia, Jinda, et al.
Published: (2024)
StableShard: Stable and Scalable Blockchain Sharding with High Concurrency via Collaborative Committees
by: Li, Mingzhe, et al.
Published: (2024)
by: Li, Mingzhe, et al.
Published: (2024)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
by: Kang, Xueze, et al.
Published: (2025)
by: Kang, Xueze, et al.
Published: (2025)
Memory and Bandwidth are All You Need for Fully Sharded Data Parallel
by: Wang, Jiangtao, et al.
Published: (2025)
by: Wang, Jiangtao, et al.
Published: (2025)
SAMM: Sharded Automated Market Maker
by: Chen, Hongyin, et al.
Published: (2024)
by: Chen, Hongyin, et al.
Published: (2024)
Fast Transaction Scheduling in Blockchain Sharding
by: Adhikari, Ramesh, et al.
Published: (2024)
by: Adhikari, Ramesh, et al.
Published: (2024)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
by: Bhatia, Nidhi, et al.
Published: (2025)
by: Bhatia, Nidhi, et al.
Published: (2025)
Self-healing Nodes with Adaptive Data-Sharding
by: Thakur, Ayush, et al.
Published: (2024)
by: Thakur, Ayush, et al.
Published: (2024)
On the Efficiency of Dynamic Transaction Scheduling in Blockchain Sharding
by: Adhikari, Ramesh, et al.
Published: (2025)
by: Adhikari, Ramesh, et al.
Published: (2025)
Sharding Distributed Databases: A Critical Review
by: Solat, Siamak
Published: (2024)
by: Solat, Siamak
Published: (2024)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
by: Ma, Tenghui, et al.
Published: (2026)
by: Ma, Tenghui, et al.
Published: (2026)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
A Hierarchical Sharded Blockchain Balancing Performance and Availability
by: Jo, Yongrae, et al.
Published: (2025)
by: Jo, Yongrae, et al.
Published: (2025)
Dynamically Sharded Ledgers on a Distributed Hash Table
by: Fink, Christoffer, et al.
Published: (2024)
by: Fink, Christoffer, et al.
Published: (2024)
Stable Blockchain Sharding under Adversarial Transaction Generation
by: Adhikari, Ramesh, et al.
Published: (2024)
by: Adhikari, Ramesh, et al.
Published: (2024)
Investigating Sharding Advancements, Methodologies, and Adoption Potential in Hedera
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
by: Ai, Xin, et al.
Published: (2024)
by: Ai, Xin, et al.
Published: (2024)
Near-Optimal Stability for Distributed Transaction Processing in Blockchain Sharding
by: Adhikari, Ramesh, et al.
Published: (2025)
by: Adhikari, Ramesh, et al.
Published: (2025)
SmartShards: Churn-Tolerant Continuously Available Distributed Ledger
by: Oglio, Joseph, et al.
Published: (2025)
by: Oglio, Joseph, et al.
Published: (2025)
EvoSort: A Genetic-Algorithm-Based Adaptive Parallel Sorting Framework for Large-Scale High Performance Computing
by: Raj, Shashank, et al.
Published: (2025)
by: Raj, Shashank, et al.
Published: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
by: Lin, Haoran, et al.
Published: (2025)
by: Lin, Haoran, et al.
Published: (2025)
Similar Items
-
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
by: Gu, Diandian, et al.
Published: (2024) -
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
by: Chen, Qiaoling, et al.
Published: (2023) -
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
by: Chen, Qiaoling, et al.
Published: (2025) -
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
by: Chen, Qiaoling, et al.
Published: (2026) -
Characterization of Large Language Model Development in the Datacenter
by: Hu, Qinghao, et al.
Published: (2024)