Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Jianli, Lin, Ziyang, Dong, Qianli, Chen, Yi, Srinivasa, Jayanth, Lee, Myungjin, Tan, Zhaowei, Lai, Fan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911315543457792
author Jin, Jianli
Lin, Ziyang
Dong, Qianli
Chen, Yi
Srinivasa, Jayanth
Lee, Myungjin
Tan, Zhaowei
Lai, Fan
author_facet Jin, Jianli
Lin, Ziyang
Dong, Qianli
Chen, Yi
Srinivasa, Jayanth
Lee, Myungjin
Tan, Zhaowei
Lai, Fan
contents With the proliferation of edge AI applications, satisfying user quality of experience (QoE) requirements, such as model inference latency, has become a first class objective, as these models operate in resource constrained settings and directly interact with users. Yet, modern AI models routinely exceed the resource capacity of individual devices, necessitating distributed execution across heterogeneous devices over variable and contention prone networks. Existing planners for hybrid (e.g., data and pipeline) parallelism largely optimize for throughput or device utilization, overlooking QoE, leading to severe resource inefficiency (e.g., unnecessary energy drain) or QoE violations under runtime dynamics. We present Dora, a framework for QoE aware hybrid parallelism in distributed edge AI training and inference. Dora jointly optimizes heterogeneous computation, contention prone networks, and multi dimensional QoE objectives via three key mechanisms: (i) a heterogeneity aware model partitioner that determines and assigns model partitions across devices, forming a compact set of QoE compliant plans; (ii) a contention aware network scheduler that further refines these candidate plans by maximizing compute communication overlap; and (iii) a runtime adapter that adaptively composes multiple plans to maximize global efficiency while respecting overall QoEs. Across representative edge deployments, including smart homes, traffic analytics, and small edge clusters, Dora achieves 1.1--6.3 times faster execution and, alternatively, reduces energy consumption by 21--82 percent, all while maintaining QoE under runtime dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10990
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
Jin, Jianli
Lin, Ziyang
Dong, Qianli
Chen, Yi
Srinivasa, Jayanth
Lee, Myungjin
Tan, Zhaowei
Lai, Fan
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
With the proliferation of edge AI applications, satisfying user quality of experience (QoE) requirements, such as model inference latency, has become a first class objective, as these models operate in resource constrained settings and directly interact with users. Yet, modern AI models routinely exceed the resource capacity of individual devices, necessitating distributed execution across heterogeneous devices over variable and contention prone networks. Existing planners for hybrid (e.g., data and pipeline) parallelism largely optimize for throughput or device utilization, overlooking QoE, leading to severe resource inefficiency (e.g., unnecessary energy drain) or QoE violations under runtime dynamics. We present Dora, a framework for QoE aware hybrid parallelism in distributed edge AI training and inference. Dora jointly optimizes heterogeneous computation, contention prone networks, and multi dimensional QoE objectives via three key mechanisms: (i) a heterogeneity aware model partitioner that determines and assigns model partitions across devices, forming a compact set of QoE compliant plans; (ii) a contention aware network scheduler that further refines these candidate plans by maximizing compute communication overlap; and (iii) a runtime adapter that adaptively composes multiple plans to maximize global efficiency while respecting overall QoEs. Across representative edge deployments, including smart homes, traffic analytics, and small edge clusters, Dora achieves 1.1--6.3 times faster execution and, alternatively, reduces energy consumption by 21--82 percent, all while maintaining QoE under runtime dynamics.
title Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2512.10990