Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhanghan, Ding, Ding, Zhu, Hang, Lin, Haibin, Panda, Aurojit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Verifiable Federated Unlearning: Framework, Challenges, and The Road Ahead
by: Nguyen, Thanh Linh, et al.
Published: (2025)
by: Nguyen, Thanh Linh, et al.
Published: (2025)
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Demystifying the Communication Characteristics for Distributed Transformer Models
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
by: Wang, Li, et al.
Published: (2024)
by: Wang, Li, et al.
Published: (2024)
ECCENTRIC: Edge-Cloud Collaboration Framework for Distributed Inference Using Knowledge Adaptation
by: Kamani, Mohammad Mahdi, et al.
Published: (2025)
by: Kamani, Mohammad Mahdi, et al.
Published: (2025)
Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
by: Chen, Daoyuan, et al.
Published: (2024)
by: Chen, Daoyuan, et al.
Published: (2024)
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
Communication-Efficient Large-Scale Distributed Deep Learning: A Comprehensive Survey
by: Liang, Feng, et al.
Published: (2024)
by: Liang, Feng, et al.
Published: (2024)
Understanding Stragglers in Large Model Training Using What-if Analysis
by: Lin, Jinkun, et al.
Published: (2025)
by: Lin, Jinkun, et al.
Published: (2025)
High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments
by: Rodriguez, Julian, et al.
Published: (2025)
by: Rodriguez, Julian, et al.
Published: (2025)
Placement Semantics for Distributed Deep Learning: A Systematic Framework for Analyzing Parallelism Strategies
by: Mehta, Deep Pankajbhai
Published: (2026)
by: Mehta, Deep Pankajbhai
Published: (2026)
Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey
by: Liang, Feng, et al.
Published: (2024)
by: Liang, Feng, et al.
Published: (2024)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
by: Qazi, Muhammad Azlan, et al.
Published: (2026)
by: Qazi, Muhammad Azlan, et al.
Published: (2026)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
by: Lu, Yunchi, et al.
Published: (2025)
by: Lu, Yunchi, et al.
Published: (2025)
Seesaw: High-throughput LLM Inference via Model Re-sharding
by: Su, Qidong, et al.
Published: (2025)
by: Su, Qidong, et al.
Published: (2025)
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
by: The AIBrix Team, et al.
Published: (2025)
by: The AIBrix Team, et al.
Published: (2025)
Distributed Inference on Mobile Edge and Cloud: A Data-Cartography based Clustering Approach
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
by: Li, Wanqian, et al.
Published: (2026)
by: Li, Wanqian, et al.
Published: (2026)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
Threats and Defenses in Federated Learning Life Cycle: A Comprehensive Survey and Challenges
by: Li, Yanli, et al.
Published: (2024)
by: Li, Yanli, et al.
Published: (2024)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
by: Bai, Fan, et al.
Published: (2026)
by: Bai, Fan, et al.
Published: (2026)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
by: Yu, Dianhai, et al.
Published: (2022)
by: Yu, Dianhai, et al.
Published: (2022)
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
by: Xu, Si, et al.
Published: (2024)
by: Xu, Si, et al.
Published: (2024)
Learning Provably Correct Distributed Protocols Without Human Knowledge
by: Hui, Yujie, et al.
Published: (2026)
by: Hui, Yujie, et al.
Published: (2026)
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
by: Gond, Raja, et al.
Published: (2026)
by: Gond, Raja, et al.
Published: (2026)
DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
by: Peng, Xian, et al.
Published: (2025)
by: Peng, Xian, et al.
Published: (2025)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
EdgeRL: Reinforcement Learning-driven Deep Learning Model Inference Optimization at Edge
by: Mounesan, Motahare, et al.
Published: (2024)
by: Mounesan, Motahare, et al.
Published: (2024)
Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization
by: Zhu, Zhanda, et al.
Published: (2025)
by: Zhu, Zhanda, et al.
Published: (2025)
Large Language Model Partitioning for Low-Latency Inference at the Edge
by: Kafetzis, Dimitrios, et al.
Published: (2025)
by: Kafetzis, Dimitrios, et al.
Published: (2025)
Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
by: An, Wei, et al.
Published: (2024)
by: An, Wei, et al.
Published: (2024)
The Case for Co-Designing Model Architectures with Hardware
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
by: He, Zifan, et al.
Published: (2026)
by: He, Zifan, et al.
Published: (2026)
Deep Reinforcement Learning for Job Scheduling and Resource Management in Cloud Computing: An Algorithm-Level Review
by: Gu, Yan, et al.
Published: (2025)
by: Gu, Yan, et al.
Published: (2025)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
by: Xu, Lang, et al.
Published: (2024)
by: Xu, Lang, et al.
Published: (2024)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
by: Chen, Yanxi, et al.
Published: (2023)
by: Chen, Yanxi, et al.
Published: (2023)
Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
Tesserae: Scalable Placement Policies for Deep Learning Workloads
by: Bian, Song, et al.
Published: (2025)
by: Bian, Song, et al.
Published: (2025)
Similar Items
-
Towards Verifiable Federated Unlearning: Framework, Challenges, and The Road Ahead
by: Nguyen, Thanh Linh, et al.
Published: (2025) -
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
by: Zhang, Chen, et al.
Published: (2026) -
Demystifying the Communication Characteristics for Distributed Transformer Models
by: Anthony, Quentin, et al.
Published: (2024) -
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
by: Wang, Li, et al.
Published: (2024) -
ECCENTRIC: Edge-Cloud Collaboration Framework for Distributed Inference Using Knowledge Adaptation
by: Kamani, Mohammad Mahdi, et al.
Published: (2025)