Inference Offloading for Cost-Sensitive Binary Classification at the Edge
Fuente:
arXiv
Saved in:
| Main Authors: | Moothedath, Vishnu Narayanan, Agarwal, Umang, N, Umeshraja, Gross, James Richard, Champati, Jaya Prakash, Moharir, Sharayu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
by: Behera, Adarsh Prasad, et al.
Published: (2024)
by: Behera, Adarsh Prasad, et al.
Published: (2024)
Cascading Bandits With Feedback
by: Prakash, R Sri, et al.
Published: (2025)
by: Prakash, R Sri, et al.
Published: (2025)
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
by: Behera, Adarsh Prasad, et al.
Published: (2025)
by: Behera, Adarsh Prasad, et al.
Published: (2025)
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
by: Behera, Adarsh Prasad, et al.
Published: (2024)
by: Behera, Adarsh Prasad, et al.
Published: (2024)
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
by: Ng, Nathan, et al.
Published: (2025)
by: Ng, Nathan, et al.
Published: (2025)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2026)
by: Yang, Zheming, et al.
Published: (2026)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
by: Cotter, Jamie, et al.
Published: (2025)
by: Cotter, Jamie, et al.
Published: (2025)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2025)
by: Yang, Zheming, et al.
Published: (2025)
Egret: Reinforcement Mechanism for Sequential Computation Offloading in Edge Computing
by: Peng, Haosong, et al.
Published: (2024)
by: Peng, Haosong, et al.
Published: (2024)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
by: Xu, Yaodan, et al.
Published: (2025)
by: Xu, Yaodan, et al.
Published: (2025)
Collaborative Inference for Large Models with Task Offloading and Early Exiting
by: Xie, Zuan, et al.
Published: (2024)
by: Xie, Zuan, et al.
Published: (2024)
FRESCO: Fast and Reliable Edge Offloading with Reputation-based Hybrid Smart Contracts
by: Zilic, Josip, et al.
Published: (2024)
by: Zilic, Josip, et al.
Published: (2024)
LLM-Enhanced Deep Reinforcement Learning for Task Offloading in Collaborative Edge Computing
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
Online Algorithms for Hierarchical Inference in Deep Learning applications at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2023)
by: Moothedath, Vishnu Narayanan, et al.
Published: (2023)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
by: Ma, Chenxiang, et al.
Published: (2025)
by: Ma, Chenxiang, et al.
Published: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
by: Wang, Zhibin, et al.
Published: (2025)
by: Wang, Zhibin, et al.
Published: (2025)
Neuro-Inspired Task Offloading in Edge-IoT Networks Using Spiking Neural Networks
by: Rossi, Fabio Diniz
Published: (2025)
by: Rossi, Fabio Diniz
Published: (2025)
Orchestrating Joint Offloading and Scheduling for Low-Latency Edge SLAM
by: Zhang, Yao, et al.
Published: (2025)
by: Zhang, Yao, et al.
Published: (2025)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
by: Lin, Shouxu, et al.
Published: (2026)
by: Lin, Shouxu, et al.
Published: (2026)
Proposal of Automatic Offloading Method in Mixed Offloading Destination Environment
by: Yamato, Yoji
Published: (2020)
by: Yamato, Yoji
Published: (2020)
Workload Distribution with Rateless Encoding: A Low-Latency Computation Offloading Method within Edge Networks
by: Guo, Zhongfu, et al.
Published: (2023)
by: Guo, Zhongfu, et al.
Published: (2023)
Energy-Efficient Joint Offloading and Resource Allocation for Deadline-Constrained Tasks in Multi-Access Edge Computing
by: Gao, Chuanchao, et al.
Published: (2025)
by: Gao, Chuanchao, et al.
Published: (2025)
A Joint Time and Energy-Efficient Federated Learning-based Computation Offloading Method for Mobile Edge Computing
by: Mukherjee, Anwesha, et al.
Published: (2024)
by: Mukherjee, Anwesha, et al.
Published: (2024)
Blockchain-Enhanced Offloading in Mobile Edge Computing: A Systematic Review and Survey of Current Trends and Future Directions
by: Moghaddasi, Komeil, et al.
Published: (2024)
by: Moghaddasi, Komeil, et al.
Published: (2024)
Edge Offloading in Smart Grid
by: Arcas, Gabriel Ioan, et al.
Published: (2024)
by: Arcas, Gabriel Ioan, et al.
Published: (2024)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
Cost-Effective Edge Data Distribution with End-To-End Delay Guarantees in Edge Computing
by: Shankar, Ravi, et al.
Published: (2025)
by: Shankar, Ravi, et al.
Published: (2025)
Plug & Offload: Transparently Offloading TCP Stack onto Off-path SmartNIC with PnO-TCP
by: Nan, Hailong, et al.
Published: (2025)
by: Nan, Hailong, et al.
Published: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
by: Jeong, Bodon, et al.
Published: (2026)
by: Jeong, Bodon, et al.
Published: (2026)
OffloadFS: Leveraging Disaggregated Storage for Computation Offloading
by: Moon, Sungho, et al.
Published: (2026)
by: Moon, Sungho, et al.
Published: (2026)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
by: Arya, Mayank, et al.
Published: (2025)
by: Arya, Mayank, et al.
Published: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
by: Rajashekar, Kolichala, et al.
Published: (2025)
by: Rajashekar, Kolichala, et al.
Published: (2025)
Environment-Aware Dynamic Pruning for Pipelined Edge Inference
by: O'Quinn, Austin, et al.
Published: (2025)
by: O'Quinn, Austin, et al.
Published: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
RIPPLE++: An Incremental Framework for Efficient GNN Inference on Evolving Graphs
by: Naman, Pranjal, et al.
Published: (2026)
by: Naman, Pranjal, et al.
Published: (2026)
A Survey of Computation Offloading with Task Types
by: Zhang, Siqi, et al.
Published: (2023)
by: Zhang, Siqi, et al.
Published: (2023)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
by: K., Prashanthi S., et al.
Published: (2023)
by: K., Prashanthi S., et al.
Published: (2023)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
by: K., Prashanthi S., et al.
Published: (2025)
by: K., Prashanthi S., et al.
Published: (2025)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
by: Khoshsirat, Aria, et al.
Published: (2024)
by: Khoshsirat, Aria, et al.
Published: (2024)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
by: Nguyen, Thanh-Tung, et al.
Published: (2025)
by: Nguyen, Thanh-Tung, et al.
Published: (2025)
Similar Items
-
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
by: Behera, Adarsh Prasad, et al.
Published: (2024) -
Cascading Bandits With Feedback
by: Prakash, R Sri, et al.
Published: (2025) -
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
by: Behera, Adarsh Prasad, et al.
Published: (2025) -
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
by: Behera, Adarsh Prasad, et al.
Published: (2024) -
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
by: Ng, Nathan, et al.
Published: (2025)