FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Hongchao, Wu, Shangyu, Kharlamova, Arina, Guan, Nan, Xue, Chun Jason |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlexBSO: Flexible Block Storage Offload for Datacenters
by: Aschenbrenner, Vojtech, et al.
Published: (2024)
by: Aschenbrenner, Vojtech, et al.
Published: (2024)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026)
by: Zou, Jing, et al.
Published: (2026)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
EvoP: Robust LLM Inference via Evolutionary Pruning
by: Wu, Shangyu, et al.
Published: (2025)
by: Wu, Shangyu, et al.
Published: (2025)
SSV: Sparse Speculative Verification for Efficient LLM Inference
by: Wang, Zhibin, et al.
Published: (2026)
by: Wang, Zhibin, et al.
Published: (2026)
HeteroPod: XPU-Accelerated Infrastructure Offloading for Commodity Cloud-Native Applications
by: Yang, Bicheng, et al.
Published: (2025)
by: Yang, Bicheng, et al.
Published: (2025)
Rethinking Inter-Process Communication with Memory Operation Offloading
by: Park, Misun, et al.
Published: (2026)
by: Park, Misun, et al.
Published: (2026)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
by: Tan, Boyu, et al.
Published: (2026)
by: Tan, Boyu, et al.
Published: (2026)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
Wave: Offloading Resource Management to SmartNIC Cores
by: Humphries, Jack Tigar, et al.
Published: (2024)
by: Humphries, Jack Tigar, et al.
Published: (2024)
Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency
by: Zhang, Zongpu, et al.
Published: (2025)
by: Zhang, Zongpu, et al.
Published: (2025)
FRAP: A Flexible Resource Accessing Protocol for Multiprocessor Real-Time Systems
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Efficient Memory Tiering in a Virtual Machine
by: Prakash, Chandra, et al.
Published: (2025)
by: Prakash, Chandra, et al.
Published: (2025)
Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
by: Xia, Tian, et al.
Published: (2026)
by: Xia, Tian, et al.
Published: (2026)
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
by: Huang, Zhengxiang, et al.
Published: (2025)
by: Huang, Zhengxiang, et al.
Published: (2025)
RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
by: Yu, Weiping, et al.
Published: (2025)
by: Yu, Weiping, et al.
Published: (2025)
Ariadne: A Hotness-Aware and Size-Adaptive Compressed Swap Technique for Fast Application Relaunch and Reduced CPU Usage on Mobile Devices
by: Liang, Yu, et al.
Published: (2025)
by: Liang, Yu, et al.
Published: (2025)
AgenTEE: Confidential LLM Agent Execution on Edge Devices
by: Abdollahi, Sina, et al.
Published: (2026)
by: Abdollahi, Sina, et al.
Published: (2026)
Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration
by: Xiang, Lingfeng, et al.
Published: (2024)
by: Xiang, Lingfeng, et al.
Published: (2024)
Towards Fully-fledged GPU Multitasking via Proactive Memory Scheduling
by: Shen, Weihang, et al.
Published: (2025)
by: Shen, Weihang, et al.
Published: (2025)
Vmem: A Lightweight Hot-Upgradable Memory Management for In-production Cloud Environment
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK
by: Zhao, Haoru, et al.
Published: (2025)
by: Zhao, Haoru, et al.
Published: (2025)
Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU
by: Sun, He, et al.
Published: (2025)
by: Sun, He, et al.
Published: (2025)
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
by: Wu, Fangnuo, et al.
Published: (2025)
by: Wu, Fangnuo, et al.
Published: (2025)
Semantic Scheduling for LLM Inference
by: Hua, Wenyue, et al.
Published: (2025)
by: Hua, Wenyue, et al.
Published: (2025)
Adaptive and Efficient Dynamic Memory Management for Hardware Enclaves
by: Dhanraj, Vijay, et al.
Published: (2025)
by: Dhanraj, Vijay, et al.
Published: (2025)
TierBPF: Page Migration Admission Control for Tiered Memory via eBPF
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Exposing Hidden Interfaces: LLM-Guided Type Inference for Reverse Engineering macOS Private Frameworks
by: Kharlamova, Arina, et al.
Published: (2026)
by: Kharlamova, Arina, et al.
Published: (2026)
uTNT: Unikernels for Efficient and Flexible Internet Probing
by: Letemple, Maxime, et al.
Published: (2024)
by: Letemple, Maxime, et al.
Published: (2024)
GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
by: Zhang, Keyao, et al.
Published: (2025)
by: Zhang, Keyao, et al.
Published: (2025)
ParaCell: Paravirtualized Secure Containers with Lightweight Intra-Container Isolation and Intent-Driven Memory Management
by: Wu, Yiyang, et al.
Published: (2026)
by: Wu, Yiyang, et al.
Published: (2026)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)
by: Song, Yixin, et al.
Published: (2023)
The First Principle of Big Memory Systems
by: Hua, Yu
Published: (2023)
by: Hua, Yu
Published: (2023)
Virtual-Memory Assisted Buffer Management In Tiered Memory
by: Rayhan, Yeasir, et al.
Published: (2026)
by: Rayhan, Yeasir, et al.
Published: (2026)
HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
by: Luo, Cheng, et al.
Published: (2025)
by: Luo, Cheng, et al.
Published: (2025)
Data-driven Software-based Power Estimation for Embedded Devices
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
by: Park, JooYoung, et al.
Published: (2026)
by: Park, JooYoung, et al.
Published: (2026)
Hybrid Adaptive Tuning for Tiered Memory Systems
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
ARMS: Adaptive and Robust Memory Tiering System
by: Yadalam, Sujay, et al.
Published: (2025)
by: Yadalam, Sujay, et al.
Published: (2025)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
Similar Items
-
FlexBSO: Flexible Block Storage Offload for Datacenters
by: Aschenbrenner, Vojtech, et al.
Published: (2024) -
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026) -
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026) -
EvoP: Robust LLM Inference via Evolutionary Pruning
by: Wu, Shangyu, et al.
Published: (2025) -
SSV: Sparse Speculative Verification for Efficient LLM Inference
by: Wang, Zhibin, et al.
Published: (2026)