Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
Fuente:
arXiv
Saved in:
| Main Authors: | Behera, Adarsh Prasad, Daubaris, Paulius, Bravo, Iñaki, Gallego, José, Morabito, Roberto, Widmer, Joerg, Champati, Jaya Prakash Varma |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
by: Behera, Adarsh Prasad, et al.
Published: (2024)
by: Behera, Adarsh Prasad, et al.
Published: (2024)
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
by: Behera, Adarsh Prasad, et al.
Published: (2025)
by: Behera, Adarsh Prasad, et al.
Published: (2025)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
2D-AoI: Age-of-Information of Distributed Sensors for Spatio-Temporal Processes
by: Fidler, Markus, et al.
Published: (2024)
by: Fidler, Markus, et al.
Published: (2024)
Online Algorithms for Hierarchical Inference in Deep Learning applications at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2023)
by: Moothedath, Vishnu Narayanan, et al.
Published: (2023)
Low-Regret and Low-Complexity Learning for Hierarchical Inference
by: Chattopadhyay, Sameep, et al.
Published: (2025)
by: Chattopadhyay, Sameep, et al.
Published: (2025)
Error Bounds for the Network Scale-Up Method
by: Díaz-Aranda, Sergio, et al.
Published: (2024)
by: Díaz-Aranda, Sergio, et al.
Published: (2024)
Pico-Cloud: Cloud Infrastructure for Tiny Edge Devices
by: Guri, Mordechai
Published: (2025)
by: Guri, Mordechai
Published: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
by: Zhang, Haolin, et al.
Published: (2025)
by: Zhang, Haolin, et al.
Published: (2025)
Inference Acceleration for Large Language Models on CPUs
by: PS, Ditto, et al.
Published: (2024)
by: PS, Ditto, et al.
Published: (2024)
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs
by: Jang, SiYoung, et al.
Published: (2025)
by: Jang, SiYoung, et al.
Published: (2025)
Future-Proofing Mobile Networks: A Digital Twin Approach to Multi-Signal Management
by: Morabito, Roberto, et al.
Published: (2024)
by: Morabito, Roberto, et al.
Published: (2024)
Sometimes Painful but Certainly Promising: Feasibility and Trade-offs of Language Model Inference at the Edge
by: Abstreiter, Maximilian, et al.
Published: (2025)
by: Abstreiter, Maximilian, et al.
Published: (2025)
Efficient Federated Finetuning of Tiny Transformers with Resource-Constrained Devices
by: Pfeiffer, Kilian, et al.
Published: (2024)
by: Pfeiffer, Kilian, et al.
Published: (2024)
Agentic TinyML for Intent-aware Handover in 6G Wireless Networks
by: Saleh, Alaa, et al.
Published: (2025)
by: Saleh, Alaa, et al.
Published: (2025)
Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices
by: Li, Yilong, et al.
Published: (2025)
by: Li, Yilong, et al.
Published: (2025)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
by: Wang, Irene, et al.
Published: (2026)
by: Wang, Irene, et al.
Published: (2026)
Where Do the Joules Go? Diagnosing Inference Energy Consumption
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
Device Sampling and Resource Optimization for Federated Learning in Cooperative Edge Networks
by: Wang, Su, et al.
Published: (2023)
by: Wang, Su, et al.
Published: (2023)
Distributed On-Device LLM Inference With Over-the-Air Computation
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Inference Load-Aware Orchestration for Hierarchical Federated Learning
by: Lackinger, Anna, et al.
Published: (2024)
by: Lackinger, Anna, et al.
Published: (2024)
Enhancing Predictive Maintenance in Mining Mobile Machinery through a TinyML-enabled Hierarchical Inference Network
by: de la Fuente, Raúl, et al.
Published: (2024)
by: de la Fuente, Raúl, et al.
Published: (2024)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
by: Liu, Kaiwei, et al.
Published: (2025)
by: Liu, Kaiwei, et al.
Published: (2025)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
by: Tayal, Mumuksh, et al.
Published: (2025)
by: Tayal, Mumuksh, et al.
Published: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
by: Chow, Will
Published: (2025)
by: Chow, Will
Published: (2025)
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
by: Yang, Xiang, et al.
Published: (2022)
by: Yang, Xiang, et al.
Published: (2022)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
by: Lin, Zheng, et al.
Published: (2024)
by: Lin, Zheng, et al.
Published: (2024)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
by: Sun, Mingyu, et al.
Published: (2025)
by: Sun, Mingyu, et al.
Published: (2025)
Minimizing Age of Detection for a Markov Source over a Lossy Channel
by: Garde, Shivang, et al.
Published: (2025)
by: Garde, Shivang, et al.
Published: (2025)
Practical Performance Guarantees for Pipelined DNN Inference
by: Archer, Aaron, et al.
Published: (2023)
by: Archer, Aaron, et al.
Published: (2023)
Device Scheduling and Assignment in Hierarchical Federated Learning for Internet of Things
by: Zhang, Tinghao, et al.
Published: (2024)
by: Zhang, Tinghao, et al.
Published: (2024)
Token Level Routing Inference System for Edge Devices
by: She, Jianshu, et al.
Published: (2025)
by: She, Jianshu, et al.
Published: (2025)
Smaller, Smarter, Closer: The Edge of Collaborative Generative AI
by: Morabito, Roberto, et al.
Published: (2025)
by: Morabito, Roberto, et al.
Published: (2025)
Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital Twin-Assisted Approach
by: Hu, Shisheng, et al.
Published: (2024)
by: Hu, Shisheng, et al.
Published: (2024)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
by: Zhang, Runhua, et al.
Published: (2025)
by: Zhang, Runhua, et al.
Published: (2025)
Adaptive Stream Processing on Edge Devices through Active Inference
by: Sedlak, Boris, et al.
Published: (2024)
by: Sedlak, Boris, et al.
Published: (2024)
Kraken: Inherently Parallel Transformers For Efficient Multi-Device Inference
by: Prabhakar, Rohan Baskar, et al.
Published: (2024)
by: Prabhakar, Rohan Baskar, et al.
Published: (2024)
Recipes for Pre-training LLMs with MXFP8
by: Mishra, Asit, et al.
Published: (2025)
by: Mishra, Asit, et al.
Published: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
by: Zhao, Xuanlei, et al.
Published: (2024)
by: Zhao, Xuanlei, et al.
Published: (2024)
Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Similar Items
-
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
by: Behera, Adarsh Prasad, et al.
Published: (2024) -
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
by: Behera, Adarsh Prasad, et al.
Published: (2025) -
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025) -
2D-AoI: Age-of-Information of Distributed Sensors for Spatio-Temporal Processes
by: Fidler, Markus, et al.
Published: (2024) -
Online Algorithms for Hierarchical Inference in Deep Learning applications at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2023)