A Scene-aware Models Adaptation Scheme for Cross-scene Online Inference on Mobile Devices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yunzhe, Zhu, Hongzi, Deng, Zhuohong, Cheng, Yunlong, Zheng, Zimu, Zhang, Liang, Chang, Shan, Guo, Minyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adacc: An Adaptive Framework Unifying Compression and Activation Recomputation for LLM Training
von: Chen, Ping, et al.
Veröffentlicht: (2025)
von: Chen, Ping, et al.
Veröffentlicht: (2025)
EchoPFL: Asynchronous Personalized Federated Learning on Mobile Devices with On-Demand Staleness Control
von: Li, Xiaochen, et al.
Veröffentlicht: (2024)
von: Li, Xiaochen, et al.
Veröffentlicht: (2024)
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
von: Yang, Xiang, et al.
Veröffentlicht: (2022)
von: Yang, Xiang, et al.
Veröffentlicht: (2022)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
von: Lin, Zheng, et al.
Veröffentlicht: (2024)
von: Lin, Zheng, et al.
Veröffentlicht: (2024)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
von: Chen, Xiaopei, et al.
Veröffentlicht: (2025)
von: Chen, Xiaopei, et al.
Veröffentlicht: (2025)
FedAPTA: Federated Multi-task Learning for Heterogeneous Devices with Adaptive Layer-wise Pruning and Task-aware Aggregation
von: Yu, Zhen, et al.
Veröffentlicht: (2025)
von: Yu, Zhen, et al.
Veröffentlicht: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
Distributed On-Device LLM Inference With Over-the-Air Computation
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
Energy-aware Incremental OTA Update for Flash-based Batteryless IoT Devices
von: Wei, Wei, et al.
Veröffentlicht: (2024)
von: Wei, Wei, et al.
Veröffentlicht: (2024)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
von: Liu, Kaiwei, et al.
Veröffentlicht: (2025)
von: Liu, Kaiwei, et al.
Veröffentlicht: (2025)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
von: Chow, Will
Veröffentlicht: (2025)
von: Chow, Will
Veröffentlicht: (2025)
Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
Aging-aware CPU Core Management for Embodied Carbon Amortization in Cloud LLM Inference
von: Hewage, Tharindu B., et al.
Veröffentlicht: (2025)
von: Hewage, Tharindu B., et al.
Veröffentlicht: (2025)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
von: Pang, Bowen, et al.
Veröffentlicht: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
von: Behera, Adarsh Prasad, et al.
Veröffentlicht: (2024)
von: Behera, Adarsh Prasad, et al.
Veröffentlicht: (2024)
KCES: A Workflow Containerization Scheduling Scheme Under Cloud-Edge Collaboration Framework
von: Shan, Chenggang, et al.
Veröffentlicht: (2024)
von: Shan, Chenggang, et al.
Veröffentlicht: (2024)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
von: Wang, Kun, et al.
Veröffentlicht: (2024)
von: Wang, Kun, et al.
Veröffentlicht: (2024)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital Twin-Assisted Approach
von: Hu, Shisheng, et al.
Veröffentlicht: (2024)
von: Hu, Shisheng, et al.
Veröffentlicht: (2024)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
von: Zhang, Runhua, et al.
Veröffentlicht: (2025)
von: Zhang, Runhua, et al.
Veröffentlicht: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
von: Gu, Jianfeng, et al.
Veröffentlicht: (2025)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Semantic-aware Token Selection and Resource Optimization for Communication-efficient Split Federated Fine-tuning in Edge Intelligence
von: Qiang, Xianke, et al.
Veröffentlicht: (2026)
von: Qiang, Xianke, et al.
Veröffentlicht: (2026)
Distributed Inference Performance Optimization for LLMs on CPUs
von: He, Pujiang, et al.
Veröffentlicht: (2024)
von: He, Pujiang, et al.
Veröffentlicht: (2024)
FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
von: Mei, Junyi, et al.
Veröffentlicht: (2024)
von: Mei, Junyi, et al.
Veröffentlicht: (2024)
Token Level Routing Inference System for Edge Devices
von: She, Jianshu, et al.
Veröffentlicht: (2025)
von: She, Jianshu, et al.
Veröffentlicht: (2025)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
von: Xia, Tian, et al.
Veröffentlicht: (2025)
von: Xia, Tian, et al.
Veröffentlicht: (2025)
Lattica: A Decentralized Cross-NAT Communication Framework for Scalable AI Inference and Training
von: Yang, Ween, et al.
Veröffentlicht: (2025)
von: Yang, Ween, et al.
Veröffentlicht: (2025)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
von: Cotter, Jamie, et al.
Veröffentlicht: (2025)
von: Cotter, Jamie, et al.
Veröffentlicht: (2025)
RingAda: Pipelining Large Model Fine-Tuning on Edge Devices with Scheduled Layer Unfreezing
von: Li, Liang, et al.
Veröffentlicht: (2025)
von: Li, Liang, et al.
Veröffentlicht: (2025)
Decaffe: DHT Tree-Based Online Federated Fake News Detection
von: Ching, Cheng-Wei, et al.
Veröffentlicht: (2023)
von: Ching, Cheng-Wei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Adacc: An Adaptive Framework Unifying Compression and Activation Recomputation for LLM Training
von: Chen, Ping, et al.
Veröffentlicht: (2025) -
EchoPFL: Asynchronous Personalized Federated Learning on Mobile Devices with On-Demand Staleness Control
von: Li, Xiaochen, et al.
Veröffentlicht: (2024) -
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
von: Yang, Xiang, et al.
Veröffentlicht: (2022) -
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
von: Lin, Zheng, et al.
Veröffentlicht: (2024) -
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)