Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Yangtao, Zhang, Lei, Wang, Qinlong, Zhi, Xiaoyun, Zhang, Xinlei, Jiang, Zhuo, Xu, Haohan, Wang, Lei, Song, Zuquan, Liu, Gaohong, Bai, Yang, Wang, Shuguang, Xiao, Wencong, Ye, Jianxi, Yu, Minlan, Xu, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Minder: Faulty Machine Detection for Large-scale Distributed Model Training
by: Deng, Yangtao, et al.
Published: (2024)
by: Deng, Yangtao, et al.
Published: (2024)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
by: He, Guoliang, et al.
Published: (2025)
by: He, Guoliang, et al.
Published: (2025)
Distributed Consensus Network: A Modularized Communication Framework and Reliability Probabilistic Analysis
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
Decouple and Decompose: Scaling Resource Allocation with DeDe
by: Xu, Zhiying, et al.
Published: (2024)
by: Xu, Zhiying, et al.
Published: (2024)
Robust LLM Training Infrastructure at ByteDance
by: Wan, Borui, et al.
Published: (2025)
by: Wan, Borui, et al.
Published: (2025)
BootSeer: Analyzing and Mitigating Initialization Bottlenecks in Large-Scale LLM Training
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
by: Zhang, Mingjun, et al.
Published: (2025)
by: Zhang, Mingjun, et al.
Published: (2025)
Exploiting Multicast for Accelerating Collective Communication
by: Xu, Chao, et al.
Published: (2026)
by: Xu, Chao, et al.
Published: (2026)
Mazzaroth: A High-Throughput DAG Consensus with State Root
by: Li, Haohan
Published: (2025)
by: Li, Haohan
Published: (2025)
SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
by: Xiong, Yifan, et al.
Published: (2024)
by: Xiong, Yifan, et al.
Published: (2024)
FedASTA: Federated adaptive spatial-temporal attention for traffic flow prediction
by: Li, Kaiyuan, et al.
Published: (2024)
by: Li, Kaiyuan, et al.
Published: (2024)
Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
by: Zhang, Xinyi, et al.
Published: (2024)
by: Zhang, Xinyi, et al.
Published: (2024)
Prioritizing Modalities: Flexible Importance Scheduling in Federated Multimodal Learning
by: Bian, Jieming, et al.
Published: (2024)
by: Bian, Jieming, et al.
Published: (2024)
EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
by: Yuan, Yitao, et al.
Published: (2026)
by: Yuan, Yitao, et al.
Published: (2026)
A Transverse-Read-assisted Fast Valid-Bits Collection in Stochastic Computing MACs for Energy-Efficient in-RTM DNNs
by: Wang, Jihe, et al.
Published: (2024)
by: Wang, Jihe, et al.
Published: (2024)
On Completely Edge-Independent Spanning Trees in Locally Twisted Cubes
by: Li, Xiaorui, et al.
Published: (2024)
by: Li, Xiaorui, et al.
Published: (2024)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
by: Bian, Zhuohang, et al.
Published: (2026)
by: Bian, Zhuohang, et al.
Published: (2026)
CAFE: Carbon-Aware Federated Learning in Geographically Distributed Data Centers
by: Bian, Jieming, et al.
Published: (2023)
by: Bian, Jieming, et al.
Published: (2023)
Understanding Stragglers in Large Model Training Using What-if Analysis
by: Lin, Jinkun, et al.
Published: (2025)
by: Lin, Jinkun, et al.
Published: (2025)
FlowMesh: A Service Fabric for Composable LLM Workflows
by: Shen, Junyi, et al.
Published: (2025)
by: Shen, Junyi, et al.
Published: (2025)
Workload composition smooths aggregate power demand while sustaining short-horizon ramps in AI data centers
by: Majumder, Subir, et al.
Published: (2026)
by: Majumder, Subir, et al.
Published: (2026)
Reliability is Blind: Collective Incentives for Decentralized Computing Marketplaces without Individual Behavior Information
by: Mont, Henry, et al.
Published: (2025)
by: Mont, Henry, et al.
Published: (2025)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
A402: Binding Cryptocurrency Payments to Service Execution for Agentic Commerce
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
EAT: QoS-Aware Edge-Collaborative AIGC Task Scheduling via Attention-Guided Diffusion Reinforcement Learning
by: Xu, Zhifei, et al.
Published: (2025)
by: Xu, Zhifei, et al.
Published: (2025)
Visualizing Distributed Traces in Aggregate
by: Samanta, Adrita, et al.
Published: (2024)
by: Samanta, Adrita, et al.
Published: (2024)
Efficient Data Distribution Estimation for Accelerated Federated Learning
by: Wang, Yuanli, et al.
Published: (2024)
by: Wang, Yuanli, et al.
Published: (2024)
Optimistic, Signature-Free Reliable Broadcast and Its Applications
by: Shrestha, Nibesh, et al.
Published: (2025)
by: Shrestha, Nibesh, et al.
Published: (2025)
Accelerating Hybrid Federated Learning Convergence under Partial Participation
by: Bian, Jieming, et al.
Published: (2023)
by: Bian, Jieming, et al.
Published: (2023)
Astra: Efficient and Money-saving Automatic Parallel Strategies Search on Heterogeneous GPUs
by: Wang, Peiran, et al.
Published: (2025)
by: Wang, Peiran, et al.
Published: (2025)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
by: Hankendi, Can, et al.
Published: (2026)
by: Hankendi, Can, et al.
Published: (2026)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding
by: Wang, Jiamin, et al.
Published: (2026)
by: Wang, Jiamin, et al.
Published: (2026)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
by: Lei, Kelun, et al.
Published: (2025)
by: Lei, Kelun, et al.
Published: (2025)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
by: Luo, Jiajun, et al.
Published: (2024)
by: Luo, Jiajun, et al.
Published: (2024)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
by: Mo, Zizhao, et al.
Published: (2025)
by: Mo, Zizhao, et al.
Published: (2025)
Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs
by: Xu, Yechen, et al.
Published: (2026)
by: Xu, Yechen, et al.
Published: (2026)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
by: Bai, Haoyu, et al.
Published: (2024)
by: Bai, Haoyu, et al.
Published: (2024)
A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Deadline-Aware Online Scheduling for LLM Fine-Tuning with Spot Market Predictions
by: Kong, Linggao, et al.
Published: (2025)
by: Kong, Linggao, et al.
Published: (2025)
Similar Items
-
Minder: Faulty Machine Detection for Large-scale Distributed Model Training
by: Deng, Yangtao, et al.
Published: (2024) -
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
by: He, Guoliang, et al.
Published: (2025) -
Distributed Consensus Network: A Modularized Communication Framework and Reliability Probabilistic Analysis
by: Li, Yuetai, et al.
Published: (2025) -
Decouple and Decompose: Scaling Resource Allocation with DeDe
by: Xu, Zhiying, et al.
Published: (2024) -
Robust LLM Training Infrastructure at ByteDance
by: Wan, Borui, et al.
Published: (2025)