Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Yishun, Zhang, Junhao, Yang, Zeyu, Armour, Wes |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
by: Liu, Ruitao, et al.
Published: (2026)
by: Liu, Ruitao, et al.
Published: (2026)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)
by: Su, Zhaoyuan, et al.
Published: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025)
by: Yarlagadda, Srihas, et al.
Published: (2025)
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
by: Wang, Yuanqing, et al.
Published: (2026)
by: Wang, Yuanqing, et al.
Published: (2026)
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
by: Wang, Zhengyang, et al.
Published: (2025)
by: Wang, Zhengyang, et al.
Published: (2025)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Photon: Federated LLM Pre-Training
by: Sani, Lorenzo, et al.
Published: (2024)
by: Sani, Lorenzo, et al.
Published: (2024)
Incentivised Orchestrated Training Architecture (IOTA): A Technical Primer for Release
by: Quinque, Felix, et al.
Published: (2025)
by: Quinque, Felix, et al.
Published: (2025)
AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
by: Bai, Huawei, et al.
Published: (2025)
by: Bai, Huawei, et al.
Published: (2025)
A Robust Power Model Training Framework for Cloud Native Runtime Energy Metric Exporter
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems
by: Xia, Yufan, et al.
Published: (2024)
by: Xia, Yufan, et al.
Published: (2024)
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
by: Hu, Tianhao, et al.
Published: (2026)
by: Hu, Tianhao, et al.
Published: (2026)
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
by: Hu, Zhengding, et al.
Published: (2026)
by: Hu, Zhengding, et al.
Published: (2026)
Converge Faster, Talk Less: Hessian-Informed Federated Zeroth-Order Optimization
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Achieving Dimension-Free Communication in Federated Learning via Zeroth-Order Optimization
by: Li, Zhe, et al.
Published: (2024)
by: Li, Zhe, et al.
Published: (2024)
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
by: Deng, Yangtao, et al.
Published: (2025)
by: Deng, Yangtao, et al.
Published: (2025)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
by: Jia, Jinda, et al.
Published: (2024)
by: Jia, Jinda, et al.
Published: (2024)
Grappa: Gradient-Only Communication for Scalable Graph Neural Network Training
by: Xu, Chongyang, et al.
Published: (2026)
by: Xu, Chongyang, et al.
Published: (2026)
A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication
by: Xia, Yufan, et al.
Published: (2026)
by: Xia, Yufan, et al.
Published: (2026)
SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly Systems
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
TrainMover: An Interruption-Resilient Runtime for ML Training
by: Lao, ChonLam, et al.
Published: (2024)
by: Lao, ChonLam, et al.
Published: (2024)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
by: Zhao, Pinxue, et al.
Published: (2024)
by: Zhao, Pinxue, et al.
Published: (2024)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025)
by: Lian, Xinyu, et al.
Published: (2025)
Optimizing Large Model Training through Overlapped Activation Recomputation
by: Chen, Ping, et al.
Published: (2024)
by: Chen, Ping, et al.
Published: (2024)
Optimizing Federated Learning for Scalable Power-demand Forecasting in Microgrids
by: Banerjee, Roopkatha, et al.
Published: (2025)
by: Banerjee, Roopkatha, et al.
Published: (2025)
GOGH: Correlation-Guided Orchestration of GPUs in Heterogeneous Clusters
by: Raeisi, Ahmad, et al.
Published: (2025)
by: Raeisi, Ahmad, et al.
Published: (2025)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
by: Ansaripour, Matin, et al.
Published: (2022)
by: Ansaripour, Matin, et al.
Published: (2022)
Understanding Silent Data Corruption in LLM Training
by: Ma, Jeffrey, et al.
Published: (2025)
by: Ma, Jeffrey, et al.
Published: (2025)
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2024)
by: Arfeen, Daiyaan, et al.
Published: (2024)
A Scalable Digital Twin Framework for Energy Optimization in Data Centers
by: Gonçalves, Raphael Hendrigo de Souza, et al.
Published: (2026)
by: Gonçalves, Raphael Hendrigo de Souza, et al.
Published: (2026)
ADAPT: A Self-Calibrating Proactive Autoscaler for Container Orchestration
by: Baghel, Himanshu Singh
Published: (2026)
by: Baghel, Himanshu Singh
Published: (2026)
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
by: Imani, HamidReza, et al.
Published: (2025)
by: Imani, HamidReza, et al.
Published: (2025)
Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge
by: Koch, Fernando, et al.
Published: (2025)
by: Koch, Fernando, et al.
Published: (2025)
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
by: Lu, Yunchi, et al.
Published: (2025)
by: Lu, Yunchi, et al.
Published: (2025)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
by: Kim, Heehoon, et al.
Published: (2026)
by: Kim, Heehoon, et al.
Published: (2026)
Unicron: Economizing Self-Healing LLM Training at Scale
by: He, Tao, et al.
Published: (2023)
by: He, Tao, et al.
Published: (2023)
OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
Similar Items
-
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
by: Liu, Ruitao, et al.
Published: (2026) -
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025) -
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
by: Yarlagadda, Srihas, et al.
Published: (2025) -
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
by: Wang, Yuanqing, et al.
Published: (2026) -
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR
by: Zhang, Yiqi, et al.
Published: (2026)