Accelerating Compound LLM Training Workloads with Maestro
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Xiulong, Chen, Hongqing, Peng, Jiaxuan, Zhou, Fan, Ruan, Zhixiang, Wang, Zekun, Zheng, Bo, Men, Rui, Wang, Haiquan, Zhang, Zhipeng, Chen, Langshi, Yuan, Man, Gao, Jiaqi, Qian, Zhengping, Lin, Junyang, Li, Yong, Lin, Wei, Wang, Junhua, Zhou, Jingren |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
by: Qiu, Zihan, et al.
Published: (2025)
by: Qiu, Zihan, et al.
Published: (2025)
Learning to Retrieve and Reason on Knowledge Graph through Active Self-Reflection
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training
by: Tang, Shengkun, et al.
Published: (2026)
by: Tang, Shengkun, et al.
Published: (2026)
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
by: Qiu, Zihan, et al.
Published: (2025)
by: Qiu, Zihan, et al.
Published: (2025)
PD‐L1 Scoring Models for Non‐Small Cell Lung Cancer in China: Current Status, AI‐Assisted Solutions and Future Perspectives
by: Ziling Huang, et al.
Published: (2025)
by: Ziling Huang, et al.
Published: (2025)
Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning
by: Zhou, Zhaomeng, et al.
Published: (2026)
by: Zhou, Zhaomeng, et al.
Published: (2026)
Semantic Communication System for Standard Knowledge in Power Iot Networks
by: Zhengping Lin, et al.
Published: (2025)
by: Zhengping Lin, et al.
Published: (2025)
Left symmetric algebras from DNA insertion
by: Yuan, Chen, et al.
Published: (2016)
by: Yuan, Chen, et al.
Published: (2016)
VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos
by: Lu, Dunjie, et al.
Published: (2025)
by: Lu, Dunjie, et al.
Published: (2025)
IoT-Brain: Grounding LLMs for Semantic-Spatial Sensor Scheduling
by: Zhou, Zhaomeng, et al.
Published: (2026)
by: Zhou, Zhaomeng, et al.
Published: (2026)
Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
by: Song, Chendong, et al.
Published: (2026)
by: Song, Chendong, et al.
Published: (2026)
TierBase: A Workload-Driven Cost-Optimized Key-Value Store
by: Shen, Zhitao, et al.
Published: (2025)
by: Shen, Zhitao, et al.
Published: (2025)
Human‐Centered Infrastructure Restoration: An Integrated Framework for Demand Estimation and Resource Allocation
by: Yudi Chen, et al.
Published: (2025)
by: Yudi Chen, et al.
Published: (2025)
Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution
by: Wang, Haiquan, et al.
Published: (2024)
by: Wang, Haiquan, et al.
Published: (2024)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
by: Wang, Tianze, et al.
Published: (2025)
by: Wang, Tianze, et al.
Published: (2025)
Group Sequence Policy Optimization
by: Zheng, Chujie, et al.
Published: (2025)
by: Zheng, Chujie, et al.
Published: (2025)
Flexural performance of composite beams with a novel laminated slab
by: Xiulong Chen, et al.
Published: (2025)
by: Xiulong Chen, et al.
Published: (2025)
Dynamic NOx emission prediction for a 660 MW super‐critical coal‐fired boiler using integrated time series network with adaptive filtering and deep learning
by: Zixuan Lin, et al.
Published: (2025)
by: Zixuan Lin, et al.
Published: (2025)
Cyclic salt‐spray study of pitting corrosion of coastal air‐conditioner aluminum fins with varying spacing, surface flatness, and placement angle
by: Congyun Lin, et al.
Published: (2025)
by: Congyun Lin, et al.
Published: (2025)
Improving the Landing Control Capability of Blended Wing Body Configuration Solar‐Powered UAVs by Using Swallow Tails and Distributed Propellers
by: Rui Wang, et al.
Published: (2025)
by: Rui Wang, et al.
Published: (2025)
Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inference
by: Yi, Ke, et al.
Published: (2024)
by: Yi, Ke, et al.
Published: (2024)
Analysis of Structural Characteristics of the HALE Joined‐Wing Configuration UAV
by: Junlei Sun, et al.
Published: (2025)
by: Junlei Sun, et al.
Published: (2025)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
by: Zhu, Botao, et al.
Published: (2025)
by: Zhu, Botao, et al.
Published: (2025)
Operator-Level Quantum Acceleration of Non-Logconcave Sampling
by: Leng, Jiaqi, et al.
Published: (2025)
by: Leng, Jiaqi, et al.
Published: (2025)
Efficient Long Context Fine-tuning with Chunk Flow
by: Yuan, Xiulong, et al.
Published: (2025)
by: Yuan, Xiulong, et al.
Published: (2025)
A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Transformer Training
by: Qiu, Zihan, et al.
Published: (2026)
by: Qiu, Zihan, et al.
Published: (2026)
Workload-Aware Incremental Reclustering in Cloud Data Warehouses
by: Liu, Yipeng, et al.
Published: (2026)
by: Liu, Yipeng, et al.
Published: (2026)
ProcessBench: Identifying Process Errors in Mathematical Reasoning
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
by: Zhang, Zhenru, et al.
Published: (2025)
by: Zhang, Zhenru, et al.
Published: (2025)
Exact Acceleration of Subgraph Graph Neural Networks by Eliminating Computation Redundancy
by: Tao, Qian, et al.
Published: (2024)
by: Tao, Qian, et al.
Published: (2024)
Hypoglycemic Effect of Ginsenoside Compound K Mediated by N‐Acetylserotonin Derived From Gut Microbiota
by: Su‐Tian‐Zi Huang, et al.
Published: (2025)
by: Su‐Tian‐Zi Huang, et al.
Published: (2025)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
by: Chen, Liang, et al.
Published: (2024)
by: Chen, Liang, et al.
Published: (2024)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
by: Hou, Yutao, et al.
Published: (2024)
by: Hou, Yutao, et al.
Published: (2024)
NiMoO 4 With High Oxidation States for Efficient Electrooxidation of Amines to Nitriles
by: Hao Chen, et al.
Published: (2026)
by: Hao Chen, et al.
Published: (2026)
SegMo: Segment-aligned Text to 3D Human Motion Generation
by: Dang, Bowen, et al.
Published: (2025)
by: Dang, Bowen, et al.
Published: (2025)
RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback
by: Shu, Junyang, et al.
Published: (2025)
by: Shu, Junyang, et al.
Published: (2025)
BladeDISC++: Memory Optimizations Based On Symbolic Shape
by: Yuan, Xiulong, et al.
Published: (2024)
by: Yuan, Xiulong, et al.
Published: (2024)
WAter: A Workload-Adaptive Knob Tuning System based on Workload Compression
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Language Confusion Gate: Language-Aware Decoding Through Model Self-Distillation
by: Zhang, Collin, et al.
Published: (2025)
by: Zhang, Collin, et al.
Published: (2025)
Similar Items
-
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
by: Qiu, Zihan, et al.
Published: (2025) -
Learning to Retrieve and Reason on Knowledge Graph through Active Self-Reflection
by: Zhang, Han, et al.
Published: (2025) -
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training
by: Tang, Shengkun, et al.
Published: (2026) -
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
by: Qiu, Zihan, et al.
Published: (2025) -
PD‐L1 Scoring Models for Non‐Small Cell Lung Cancer in China: Current Status, AI‐Assisted Solutions and Future Perspectives
by: Ziling Huang, et al.
Published: (2025)