Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
Fuente:
arXiv
Saved in:
| Main Authors: | Behera, Adarsh Prasad, Morabito, Roberto, Widmer, Joerg, Champati, Jaya Prakash |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
by: Behera, Adarsh Prasad, et al.
Published: (2024)
by: Behera, Adarsh Prasad, et al.
Published: (2024)
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
by: Behera, Adarsh Prasad, et al.
Published: (2025)
by: Behera, Adarsh Prasad, et al.
Published: (2025)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
by: Sun, Mingyu, et al.
Published: (2025)
by: Sun, Mingyu, et al.
Published: (2025)
Distributed Hierarchical Machine Learning for Joint Resource Allocation and Slice Selection in In-Network Edge Systems
by: Rashid, Sulaiman Muhammad, et al.
Published: (2025)
by: Rashid, Sulaiman Muhammad, et al.
Published: (2025)
Device Sampling and Resource Optimization for Federated Learning in Cooperative Edge Networks
by: Wang, Su, et al.
Published: (2023)
by: Wang, Su, et al.
Published: (2023)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
by: Zhao, Xuanlei, et al.
Published: (2024)
by: Zhao, Xuanlei, et al.
Published: (2024)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
by: Tayal, Mumuksh, et al.
Published: (2025)
by: Tayal, Mumuksh, et al.
Published: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
by: Chow, Will
Published: (2025)
by: Chow, Will
Published: (2025)
FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
by: Wu, Li, et al.
Published: (2025)
by: Wu, Li, et al.
Published: (2025)
Accelerating Local LLMs on Resource-Constrained Edge Devices via Distributed Prompt Caching
by: Matsutani, Hiroki, et al.
Published: (2026)
by: Matsutani, Hiroki, et al.
Published: (2026)
Distributed Resource Selection for Self-Organising Cloud-Edge Systems
by: Renau, Quentin, et al.
Published: (2025)
by: Renau, Quentin, et al.
Published: (2025)
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs
by: Jang, SiYoung, et al.
Published: (2025)
by: Jang, SiYoung, et al.
Published: (2025)
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
by: Sheng, Jinhao, et al.
Published: (2025)
by: Sheng, Jinhao, et al.
Published: (2025)
HiRL: Hierarchical Reinforcement Learning for Coordinated Resource Management in Heterogeneous Edge Computing
by: Zhu, Jianyong, et al.
Published: (2026)
by: Zhu, Jianyong, et al.
Published: (2026)
Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital Twin-Assisted Approach
by: Hu, Shisheng, et al.
Published: (2024)
by: Hu, Shisheng, et al.
Published: (2024)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
by: Zhang, Runhua, et al.
Published: (2025)
by: Zhang, Runhua, et al.
Published: (2025)
Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices
by: Shen, Tao, et al.
Published: (2025)
by: Shen, Tao, et al.
Published: (2025)
Sometimes Painful but Certainly Promising: Feasibility and Trade-offs of Language Model Inference at the Edge
by: Abstreiter, Maximilian, et al.
Published: (2025)
by: Abstreiter, Maximilian, et al.
Published: (2025)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
by: Dong, Jiangwen, et al.
Published: (2025)
by: Dong, Jiangwen, et al.
Published: (2025)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
by: Wolfrath, Joel, et al.
Published: (2025)
by: Wolfrath, Joel, et al.
Published: (2025)
Energy-Efficient Joint Offloading and Resource Allocation for Deadline-Constrained Tasks in Multi-Access Edge Computing
by: Gao, Chuanchao, et al.
Published: (2025)
by: Gao, Chuanchao, et al.
Published: (2025)
Improved Methods of Task Assignment and Resource Allocation with Preemption in Edge Computing Systems
by: Rublein, Caroline, et al.
Published: (2024)
by: Rublein, Caroline, et al.
Published: (2024)
Token Level Routing Inference System for Edge Devices
by: She, Jianshu, et al.
Published: (2025)
by: She, Jianshu, et al.
Published: (2025)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
by: Ng, Nathan, et al.
Published: (2026)
by: Ng, Nathan, et al.
Published: (2026)
GenAI at the Edge: Comprehensive Survey on Empowering Edge Devices
by: Navardi, Mozhgan, et al.
Published: (2025)
by: Navardi, Mozhgan, et al.
Published: (2025)
SlimEdge: Performance and Device Aware Distributed DNN Deployment on Resource-Constrained Edge Hardware
by: Kumar, Mahadev Sunil, et al.
Published: (2025)
by: Kumar, Mahadev Sunil, et al.
Published: (2025)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
by: Cotter, Jamie, et al.
Published: (2025)
by: Cotter, Jamie, et al.
Published: (2025)
AIGC-assisted Federated Learning for Vehicular Edge Intelligence: Vehicle Selection, Resource Allocation and Model Augmentation
by: Qiang, Xianke, et al.
Published: (2025)
by: Qiang, Xianke, et al.
Published: (2025)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
Semantic-aware Token Selection and Resource Optimization for Communication-efficient Split Federated Fine-tuning in Edge Intelligence
by: Qiang, Xianke, et al.
Published: (2026)
by: Qiang, Xianke, et al.
Published: (2026)
Container Data Item: An Abstract Datatype for Efficient Container-based Edge Computing
by: Rahman, Md Rezwanur, et al.
Published: (2024)
by: Rahman, Md Rezwanur, et al.
Published: (2024)
Inference Acceleration for Large Language Models on CPUs
by: PS, Ditto, et al.
Published: (2024)
by: PS, Ditto, et al.
Published: (2024)
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
by: Hong, Guihang, et al.
Published: (2025)
by: Hong, Guihang, et al.
Published: (2025)
Adaptive Stream Processing on Edge Devices through Active Inference
by: Sedlak, Boris, et al.
Published: (2024)
by: Sedlak, Boris, et al.
Published: (2024)
Smaller, Smarter, Closer: The Edge of Collaborative Generative AI
by: Morabito, Roberto, et al.
Published: (2025)
by: Morabito, Roberto, et al.
Published: (2025)
ADApt: Edge Device Anomaly Detection and Microservice Replica Prediction
by: Mehran, Narges, et al.
Published: (2025)
by: Mehran, Narges, et al.
Published: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
by: Arya, Mayank, et al.
Published: (2025)
by: Arya, Mayank, et al.
Published: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
by: Rajashekar, Kolichala, et al.
Published: (2025)
by: Rajashekar, Kolichala, et al.
Published: (2025)
Environment-Aware Dynamic Pruning for Pipelined Edge Inference
by: O'Quinn, Austin, et al.
Published: (2025)
by: O'Quinn, Austin, et al.
Published: (2025)
Similar Items
-
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
by: Behera, Adarsh Prasad, et al.
Published: (2024) -
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
by: Behera, Adarsh Prasad, et al.
Published: (2025) -
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025) -
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
by: Sun, Mingyu, et al.
Published: (2025) -
Distributed Hierarchical Machine Learning for Joint Resource Allocation and Slice Selection in In-Network Edge Systems
by: Rashid, Sulaiman Muhammad, et al.
Published: (2025)