Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Jianli, Lin, Ziyang, Dong, Qianli, Chen, Yi, Srinivasa, Jayanth, Lee, Myungjin, Tan, Zhaowei, Lai, Fan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QoE-oriented Dependent Task Scheduling under Multi-dimensional QoS Constraints over Distributed Networks
by: Fan, Xuwei, et al.
Published: (2023)
by: Fan, Xuwei, et al.
Published: (2023)
Software-Defined Agentic Serving
by: Agarwal, Saurabh, et al.
Published: (2026)
by: Agarwal, Saurabh, et al.
Published: (2026)
Nalar: An agent serving framework
by: Laju, Marco, et al.
Published: (2026)
by: Laju, Marco, et al.
Published: (2026)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
by: Xia, Mengchun, et al.
Published: (2026)
by: Xia, Mengchun, et al.
Published: (2026)
QONNECT: A QoS-Aware Orchestration System for Distributed Kubernetes Clusters
by: Aslan, Haci Ismail, et al.
Published: (2025)
by: Aslan, Haci Ismail, et al.
Published: (2025)
Mercury: QoS-Aware Tiered Memory System
by: Lu, Jiaheng, et al.
Published: (2024)
by: Lu, Jiaheng, et al.
Published: (2024)
EAT: QoS-Aware Edge-Collaborative AIGC Task Scheduling via Attention-Guided Diffusion Reinforcement Learning
by: Xu, Zhifei, et al.
Published: (2025)
by: Xu, Zhifei, et al.
Published: (2025)
QoS Aware Mixed-Criticality Task Scheduling in Vehicular Edge Cloud System
by: Sarkar, Suvarthi, et al.
Published: (2024)
by: Sarkar, Suvarthi, et al.
Published: (2024)
MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
by: Zhou, Bowen, et al.
Published: (2026)
by: Zhou, Bowen, et al.
Published: (2026)
Enabling Elastic Model Serving with MultiWorld
by: Lee, Myungjin, et al.
Published: (2024)
by: Lee, Myungjin, et al.
Published: (2024)
QEdgeProxy: QoS-Aware Load Balancing for IoT Services in the Computing Continuum
by: Čilić, Ivan, et al.
Published: (2024)
by: Čilić, Ivan, et al.
Published: (2024)
Hybrid-Parallel: Achieving High Performance and Energy Efficient Distributed Inference on Robots
by: Sun, Zekai, et al.
Published: (2024)
by: Sun, Zekai, et al.
Published: (2024)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
by: Zhang, Han, et al.
Published: (2026)
by: Zhang, Han, et al.
Published: (2026)
Edge Intelligence in Satellite-Terrestrial Networks with Hybrid Quantum Computing
by: Huang, Siyue, et al.
Published: (2024)
by: Huang, Siyue, et al.
Published: (2024)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
by: Nian, Sean, et al.
Published: (2026)
by: Nian, Sean, et al.
Published: (2026)
S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge
by: Yoon, JinYi, et al.
Published: (2025)
by: Yoon, JinYi, et al.
Published: (2025)
HyperParallel: A Supernode-Affinity AI Framework
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
by: Rashid, Md Hasanur, et al.
Published: (2026)
by: Rashid, Md Hasanur, et al.
Published: (2026)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Resource-efficient Parallel Split Learning in Heterogeneous Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
Squeezing Edge Performance: A Sensitivity-Aware Container Management for Heterogeneous Tasks
by: Zhang, Yongmin, et al.
Published: (2025)
by: Zhang, Yongmin, et al.
Published: (2025)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
by: Dong, Jiangwen, et al.
Published: (2025)
by: Dong, Jiangwen, et al.
Published: (2025)
Distributed Edge Analytics in Edge-Fog-Cloud Continuum
by: Srirama, Satish Narayana
Published: (2024)
by: Srirama, Satish Narayana
Published: (2024)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
by: Cao, Jiahe, et al.
Published: (2026)
by: Cao, Jiahe, et al.
Published: (2026)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
by: Han, Yu, et al.
Published: (2025)
by: Han, Yu, et al.
Published: (2025)
Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services
by: Liu, Jiachen, et al.
Published: (2024)
by: Liu, Jiachen, et al.
Published: (2024)
pdGRASS: A Fast Parallel Density-Aware Algorithm for Graph Spectral Sparsification
by: Zhao, Tiancheng, et al.
Published: (2025)
by: Zhao, Tiancheng, et al.
Published: (2025)
On Harnessing Idle Compute at the Edge for Foundation Model Training
by: Xue, Leyang, et al.
Published: (2025)
by: Xue, Leyang, et al.
Published: (2025)
QECO: A QoE-Oriented Computation Offloading Algorithm based on Deep Reinforcement Learning for Mobile Edge Computing
by: Rahmaty, Iman, et al.
Published: (2023)
by: Rahmaty, Iman, et al.
Published: (2023)
CaPGNN: Optimizing Parallel Graph Neural Network Training with Joint Caching and Resource-Aware Graph Partitioning
by: Song, Xianfeng, et al.
Published: (2025)
by: Song, Xianfeng, et al.
Published: (2025)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
by: Kang, Xueze, et al.
Published: (2025)
by: Kang, Xueze, et al.
Published: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
by: Luo, Ziyue, et al.
Published: (2025)
by: Luo, Ziyue, et al.
Published: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
by: Wei, Jinhui, et al.
Published: (2025)
by: Wei, Jinhui, et al.
Published: (2025)
Communication-Computation Pipeline Parallel Split Learning over Wireless Edge Networks
by: Liu, Chenyu, et al.
Published: (2025)
by: Liu, Chenyu, et al.
Published: (2025)
EdgeFaaS: A Function-based Framework for Edge Computing
by: Jin, Runyu, et al.
Published: (2022)
by: Jin, Runyu, et al.
Published: (2022)
Parallel Spawning Strategies for Dynamic-Aware MPI Applications
by: Martín-Álvarez, Iker, et al.
Published: (2025)
by: Martín-Álvarez, Iker, et al.
Published: (2025)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
by: Zhang, Yuning, et al.
Published: (2025)
by: Zhang, Yuning, et al.
Published: (2025)
GenAI at the Edge: Comprehensive Survey on Empowering Edge Devices
by: Navardi, Mozhgan, et al.
Published: (2025)
by: Navardi, Mozhgan, et al.
Published: (2025)
Similar Items
-
QoE-oriented Dependent Task Scheduling under Multi-dimensional QoS Constraints over Distributed Networks
by: Fan, Xuwei, et al.
Published: (2023) -
Software-Defined Agentic Serving
by: Agarwal, Saurabh, et al.
Published: (2026) -
Nalar: An agent serving framework
by: Laju, Marco, et al.
Published: (2026) -
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
by: Xia, Mengchun, et al.
Published: (2026) -
QONNECT: A QoS-Aware Orchestration System for Distributed Kubernetes Clusters
by: Aslan, Haci Ismail, et al.
Published: (2025)