PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Yixin, Mi, Zeyu, Xie, Haotong, Chen, Haibo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PowerInfer-2: Fast Large Language Model Inference on a Smartphone
by: Xue, Zhenliang, et al.
Published: (2024)
by: Xue, Zhenliang, et al.
Published: (2024)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024)
by: Prabhu, Ramya, et al.
Published: (2024)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
by: Wang, Yuan, et al.
Published: (2025)
by: Wang, Yuan, et al.
Published: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
Accelerated Training on Low-Power Edge Devices
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025)
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024)
by: Kamahori, Keisuke, et al.
Published: (2024)
Towards Fully-fledged GPU Multitasking via Proactive Memory Scheduling
by: Shen, Weihang, et al.
Published: (2025)
by: Shen, Weihang, et al.
Published: (2025)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
by: Siavashi, Mohammad, et al.
Published: (2026)
by: Siavashi, Mohammad, et al.
Published: (2026)
TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
Puzzle: Scheduling Multiple Deep Learning Models on Mobile Device with Heterogeneous Processors
by: Kang, Duseok, et al.
Published: (2025)
by: Kang, Duseok, et al.
Published: (2025)
Trustworthy and Controllable Professional Knowledge Utilization in Large Language Models with TEE-GPU Execution
by: Cai, Yifeng, et al.
Published: (2025)
by: Cai, Yifeng, et al.
Published: (2025)
MaLV-OS: Rethinking the Operating System Architecture for Machine Learning in Virtualized Clouds
by: Bitchebe, Stella, et al.
Published: (2025)
by: Bitchebe, Stella, et al.
Published: (2025)
Crash-Consistent Checkpointing for AI Training on macOS/APFS
by: Jeon, Juha
Published: (2025)
by: Jeon, Juha
Published: (2025)
Herding LLaMaS: Using LLMs as an OS Module
by: Kamath, Aditya K, et al.
Published: (2024)
by: Kamath, Aditya K, et al.
Published: (2024)
Reinforcement Learning for Dynamic Memory Allocation
by: Lim, Arisrei, et al.
Published: (2024)
by: Lim, Arisrei, et al.
Published: (2024)
Machine Learning (ML) library in Linux kernel
by: Dubeyko, Viacheslav
Published: (2026)
by: Dubeyko, Viacheslav
Published: (2026)
Energy-Efficient Computation with DVFS using Deep Reinforcement Learning for Multi-Task Systems in Edge Computing
by: Li, Xinyi, et al.
Published: (2024)
by: Li, Xinyi, et al.
Published: (2024)
LithOS: An Operating System for Efficient Machine Learning on GPUs
by: Coppock, Patrick H., et al.
Published: (2025)
by: Coppock, Patrick H., et al.
Published: (2025)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
by: Xu, Yuhang, et al.
Published: (2025)
by: Xu, Yuhang, et al.
Published: (2025)
EROICA: Online Performance Troubleshooting for Large-scale Model Training
by: Guan, Yu, et al.
Published: (2025)
by: Guan, Yu, et al.
Published: (2025)
Characterizing Network Requirements for GPU API Remoting in AI Applications
by: Wang, Tianxia, et al.
Published: (2024)
by: Wang, Tianxia, et al.
Published: (2024)
When eBPF Meets Machine Learning: On-the-fly OS Kernel Compartmentalization
by: Wang, Zicheng, et al.
Published: (2024)
by: Wang, Zicheng, et al.
Published: (2024)
Bauplan: zero-copy, scale-up FaaS for data pipelines
by: Tagliabue, Jacopo, et al.
Published: (2024)
by: Tagliabue, Jacopo, et al.
Published: (2024)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
by: Zhang, Dingyan, et al.
Published: (2024)
by: Zhang, Dingyan, et al.
Published: (2024)
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
Performance Isolation and Semantic Determinism in Efficient GPU Spatial Sharing
by: Yang, Zhenyuan, et al.
Published: (2026)
by: Yang, Zhenyuan, et al.
Published: (2026)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
Efficient Function-as-a-Service for Large Language Models with TIDAL
by: Cui, Weihao, et al.
Published: (2025)
by: Cui, Weihao, et al.
Published: (2025)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
by: Abhyankar, Reyna, et al.
Published: (2025)
by: Abhyankar, Reyna, et al.
Published: (2025)
An Integrated Artificial Intelligence Operating System for Advanced Low-Altitude Aviation Applications
by: Tan, Minzhe, et al.
Published: (2024)
by: Tan, Minzhe, et al.
Published: (2024)
Semantic Scheduling for LLM Inference
by: Hua, Wenyue, et al.
Published: (2025)
by: Hua, Wenyue, et al.
Published: (2025)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
by: Desai, Omkar, et al.
Published: (2025)
by: Desai, Omkar, et al.
Published: (2025)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
Enhancing Battery Storage Energy Arbitrage with Deep Reinforcement Learning and Time-Series Forecasting
by: Sage, Manuel, et al.
Published: (2024)
by: Sage, Manuel, et al.
Published: (2024)
Generative Profiling for Soft Real-Time Systems and its Applications to Resource Allocation
by: Bondar, Georgiy A., et al.
Published: (2026)
by: Bondar, Georgiy A., et al.
Published: (2026)
Similar Items
-
PowerInfer-2: Fast Large Language Model Inference on a Smartphone
by: Xue, Zhenliang, et al.
Published: (2024) -
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026) -
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024) -
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025) -
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
by: Wang, Yuan, et al.
Published: (2025)