LLM as a System Service on Mobile Devices
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Wangsong, Xu, Mengwei, Li, Yuanchun, Liu, Xuanzhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Elastic On-Device LLM Service
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
by: Yin, Wangsong, et al.
Published: (2025)
by: Yin, Wangsong, et al.
Published: (2025)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency
by: Zhang, Zongpu, et al.
Published: (2025)
by: Zhang, Zongpu, et al.
Published: (2025)
Efficient Function-as-a-Service for Large Language Models with TIDAL
by: Cui, Weihao, et al.
Published: (2025)
by: Cui, Weihao, et al.
Published: (2025)
Puzzle: Scheduling Multiple Deep Learning Models on Mobile Device with Heterogeneous Processors
by: Kang, Duseok, et al.
Published: (2025)
by: Kang, Duseok, et al.
Published: (2025)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
by: Tan, Boyu, et al.
Published: (2026)
by: Tan, Boyu, et al.
Published: (2026)
Data-driven Software-based Power Estimation for Embedded Devices
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
CARTOS: A Charging-Aware Real-Time Operating System for Intermittent Batteryless Devices
by: Karimi, Mohsen, et al.
Published: (2023)
by: Karimi, Mohsen, et al.
Published: (2023)
AppFlow: Memory Scheduling for Cold Launch of Large Apps on Mobile and Vehicle Systems
by: Li, Xiaochen, et al.
Published: (2026)
by: Li, Xiaochen, et al.
Published: (2026)
AgenTEE: Confidential LLM Agent Execution on Edge Devices
by: Abdollahi, Sina, et al.
Published: (2026)
by: Abdollahi, Sina, et al.
Published: (2026)
AgentRM: An OS-Inspired Resource Manager for LLM Agent Systems
by: She, Jianshu
Published: (2026)
by: She, Jianshu
Published: (2026)
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
by: Huang, Zhengxiang, et al.
Published: (2025)
by: Huang, Zhengxiang, et al.
Published: (2025)
Ariadne: A Hotness-Aware and Size-Adaptive Compressed Swap Technique for Fast Application Relaunch and Reduced CPU Usage on Mobile Devices
by: Liang, Yu, et al.
Published: (2025)
by: Liang, Yu, et al.
Published: (2025)
EdgeFlow: Fast Cold Starts for LLMs on Mobile Devices
by: Yan, Yongsheng, et al.
Published: (2026)
by: Yan, Yongsheng, et al.
Published: (2026)
Principled Performance Tunability in Operating System Kernels
by: Chen, Zhongjie, et al.
Published: (2025)
by: Chen, Zhongjie, et al.
Published: (2025)
Hardware Memory Management for Future Mobile Hybrid Memory Systems
by: Wen, Fei, et al.
Published: (2020)
by: Wen, Fei, et al.
Published: (2020)
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
by: Du, Hongchao, et al.
Published: (2025)
by: Du, Hongchao, et al.
Published: (2025)
ROSfs: A User-Level File System for ROS
by: Xu, Zijun, et al.
Published: (2024)
by: Xu, Zijun, et al.
Published: (2024)
Hybrid Adaptive Tuning for Tiered Memory Systems
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
by: Xia, Tian, et al.
Published: (2026)
by: Xia, Tian, et al.
Published: (2026)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
by: Zhang, Keyao, et al.
Published: (2025)
by: Zhang, Keyao, et al.
Published: (2025)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
AIOS: LLM Agent Operating System
by: Mei, Kai, et al.
Published: (2024)
by: Mei, Kai, et al.
Published: (2024)
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
by: Cui, Weihao, et al.
Published: (2025)
by: Cui, Weihao, et al.
Published: (2025)
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
by: Kang, Hao, et al.
Published: (2026)
by: Kang, Hao, et al.
Published: (2026)
Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
by: Li, Ruihao, et al.
Published: (2025)
by: Li, Ruihao, et al.
Published: (2025)
FRAP: A Flexible Resource Accessing Protocol for Multiprocessor Real-Time Systems
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
by: Liu, Jinshu, et al.
Published: (2024)
by: Liu, Jinshu, et al.
Published: (2024)
Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects
by: Ruzhanskaia, Anastasiia, et al.
Published: (2024)
by: Ruzhanskaia, Anastasiia, et al.
Published: (2024)
SSV: Sparse Speculative Verification for Efficient LLM Inference
by: Wang, Zhibin, et al.
Published: (2026)
by: Wang, Zhibin, et al.
Published: (2026)
DFUSE: Strongly Consistent Write-Back Kernel Caching for Distributed Userspace File Systems
by: Li, Haoyu, et al.
Published: (2025)
by: Li, Haoyu, et al.
Published: (2025)
ByteFS: System Support for (CXL-based) Memory-Semantic Solid-State Drives
by: Li, Shaobo, et al.
Published: (2025)
by: Li, Shaobo, et al.
Published: (2025)
TenonOS: A Self-Generating LibOS-on-LibOS Framework for Time-Critical Embedded Operating Systems
by: Zhao, Xinkui, et al.
Published: (2025)
by: Zhao, Xinkui, et al.
Published: (2025)
ConZone+: Practical Zoned Flash Storage Emulation for Consumer Devices
by: Yu, Dingcui, et al.
Published: (2025)
by: Yu, Dingcui, et al.
Published: (2025)
Columbo: Low Level End-to-End System Traces through Modular Full-System Simulation
by: Görgen, Jakob, et al.
Published: (2024)
by: Görgen, Jakob, et al.
Published: (2024)
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
by: Luo, Shutian, et al.
Published: (2026)
by: Luo, Shutian, et al.
Published: (2026)
Boosting File Systems Elegantly: A Transparent NVM Write-ahead Log for Disk File Systems
by: Wang, Guoyu, et al.
Published: (2024)
by: Wang, Guoyu, et al.
Published: (2024)
Accelerated Training on Low-Power Edge Devices
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025)
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025)
Similar Items
-
Elastic On-Device LLM Service
by: Yin, Wangsong, et al.
Published: (2024) -
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
by: Yin, Wangsong, et al.
Published: (2025) -
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026) -
Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency
by: Zhang, Zongpu, et al.
Published: (2025) -
Efficient Function-as-a-Service for Large Language Models with TIDAL
by: Cui, Weihao, et al.
Published: (2025)