DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Xian, Wu, Xin, Xu, Lianming, Wang, Li, Fei, Aiguo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
von: Wang, Li, et al.
Veröffentlicht: (2024)
von: Wang, Li, et al.
Veröffentlicht: (2024)
Distributed Inference on Mobile Edge and Cloud: An Early Exit based Clustering Approach
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
Early-Exit meets Model-Distributed Inference at Edge Networks
von: Colocrese, Marco, et al.
Veröffentlicht: (2024)
von: Colocrese, Marco, et al.
Veröffentlicht: (2024)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
von: Chen, Yanxi, et al.
Veröffentlicht: (2023)
von: Chen, Yanxi, et al.
Veröffentlicht: (2023)
Federated Learning for Collaborative Inference Systems: The Case of Early Exit Networks
von: Kaplan, Caelin, et al.
Veröffentlicht: (2024)
von: Kaplan, Caelin, et al.
Veröffentlicht: (2024)
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
ECCENTRIC: Edge-Cloud Collaboration Framework for Distributed Inference Using Knowledge Adaptation
von: Kamani, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Kamani, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
Distributed Inference on Mobile Edge and Cloud: A Data-Cartography based Clustering Approach
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
von: Wang, Zhanghan, et al.
Veröffentlicht: (2025)
von: Wang, Zhanghan, et al.
Veröffentlicht: (2025)
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
von: Zheng, Peirong, et al.
Veröffentlicht: (2026)
von: Zheng, Peirong, et al.
Veröffentlicht: (2026)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
Distributed Graph Neural Network Inference With Just-In-Time Compilation For Industry-Scale Graphs
von: Wu, Xiabao, et al.
Veröffentlicht: (2025)
von: Wu, Xiabao, et al.
Veröffentlicht: (2025)
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
von: Han, Xueyuan, et al.
Veröffentlicht: (2024)
von: Han, Xueyuan, et al.
Veröffentlicht: (2024)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
von: Deng, Xiumei, et al.
Veröffentlicht: (2025)
von: Deng, Xiumei, et al.
Veröffentlicht: (2025)
Large Language Model Partitioning for Low-Latency Inference at the Edge
von: Kafetzis, Dimitrios, et al.
Veröffentlicht: (2025)
von: Kafetzis, Dimitrios, et al.
Veröffentlicht: (2025)
Acceleration for Deep Reinforcement Learning using Parallel and Distributed Computing: A Survey
von: Liu, Zhihong, et al.
Veröffentlicht: (2024)
von: Liu, Zhihong, et al.
Veröffentlicht: (2024)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
Quality Scalable Quantization Methodology for Deep Learning on Edge
von: Khaliq, Salman Abdul, et al.
Veröffentlicht: (2024)
von: Khaliq, Salman Abdul, et al.
Veröffentlicht: (2024)
EdgeRL: Reinforcement Learning-driven Deep Learning Model Inference Optimization at Edge
von: Mounesan, Motahare, et al.
Veröffentlicht: (2024)
von: Mounesan, Motahare, et al.
Veröffentlicht: (2024)
Hardware Utilization and Inference Performance of Edge Object Detection Under Fault Injection
von: Pasandideh, Faezeh, et al.
Veröffentlicht: (2026)
von: Pasandideh, Faezeh, et al.
Veröffentlicht: (2026)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
Distributed Neural Representation for Reactive in situ Visualization
von: Wu, Qi, et al.
Veröffentlicht: (2023)
von: Wu, Qi, et al.
Veröffentlicht: (2023)
FedDCT: A Dynamic Cross-Tier Federated Learning Framework in Wireless Networks
von: Xian, Youquan, et al.
Veröffentlicht: (2023)
von: Xian, Youquan, et al.
Veröffentlicht: (2023)
DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72
von: Li, Wanqian, et al.
Veröffentlicht: (2026)
von: Li, Wanqian, et al.
Veröffentlicht: (2026)
Design and Optimization of Hierarchical Gradient Coding for Distributed Learning at Edge Devices
von: Tang, Weiheng, et al.
Veröffentlicht: (2024)
von: Tang, Weiheng, et al.
Veröffentlicht: (2024)
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
von: Jin, Jianli, et al.
Veröffentlicht: (2025)
von: Jin, Jianli, et al.
Veröffentlicht: (2025)
Elastic On-Device LLM Service
von: Yin, Wangsong, et al.
Veröffentlicht: (2024)
von: Yin, Wangsong, et al.
Veröffentlicht: (2024)
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
von: Wang, Kun, et al.
Veröffentlicht: (2024)
von: Wang, Kun, et al.
Veröffentlicht: (2024)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
von: Yu, Dianhai, et al.
Veröffentlicht: (2022)
von: Yu, Dianhai, et al.
Veröffentlicht: (2022)
DGRAG: Distributed Graph-based Retrieval-Augmented Generation in Edge-Cloud Systems
von: Zhou, Wenqing, et al.
Veröffentlicht: (2025)
von: Zhou, Wenqing, et al.
Veröffentlicht: (2025)
HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
von: Taufique, Zain, et al.
Veröffentlicht: (2024)
von: Taufique, Zain, et al.
Veröffentlicht: (2024)
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
A Model Aware AIGC Task Offloading Algorithm in IIoT Edge Computing
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
von: Zhang, Zongshun, et al.
Veröffentlicht: (2025)
von: Zhang, Zongshun, et al.
Veröffentlicht: (2025)
PIPO: Pipelined Offloading for Efficient Inference on Consumer Devices
von: Liu, Yangyijian, et al.
Veröffentlicht: (2025)
von: Liu, Yangyijian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
von: Wang, Li, et al.
Veröffentlicht: (2024) -
Distributed Inference on Mobile Edge and Cloud: An Early Exit based Clustering Approach
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024) -
Early-Exit meets Model-Distributed Inference at Edge Networks
von: Colocrese, Marco, et al.
Veröffentlicht: (2024) -
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025) -
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
von: Chen, Yanxi, et al.
Veröffentlicht: (2023)