Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lee, Sanghyeon, Kim, Hongbeen, Hwang, Soojin, Heo, Guseul, Noh, Minwoo, Huh, Jaehyuk |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
par: Cho, Jaehong, et autres
Publié: (2024)
par: Cho, Jaehong, et autres
Publié: (2024)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
par: Cho, Jaehong, et autres
Publié: (2026)
par: Cho, Jaehong, et autres
Publié: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
par: Yoon, Dongha, et autres
Publié: (2025)
par: Yoon, Dongha, et autres
Publié: (2025)
Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
par: Hwang, Jinwoo, et autres
Publié: (2025)
par: Hwang, Jinwoo, et autres
Publié: (2025)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
par: Nian, Sean, et autres
Publié: (2026)
par: Nian, Sean, et autres
Publié: (2026)
FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
par: Ma, Jian, et autres
Publié: (2025)
par: Ma, Jian, et autres
Publié: (2025)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
par: Dong, Jiangwen, et autres
Publié: (2025)
par: Dong, Jiangwen, et autres
Publié: (2025)
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
par: Gossman, Mikaila J., et autres
Publié: (2025)
par: Gossman, Mikaila J., et autres
Publié: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
par: Wang, Yuxin, et autres
Publié: (2023)
par: Wang, Yuxin, et autres
Publié: (2023)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
par: Zhang, Yuning, et autres
Publié: (2025)
par: Zhang, Yuning, et autres
Publié: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
par: Lin, Mao, et autres
Publié: (2026)
par: Lin, Mao, et autres
Publié: (2026)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
par: Kim, Kihyun, et autres
Publié: (2025)
par: Kim, Kihyun, et autres
Publié: (2025)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
par: Nicolae, Radu, et autres
Publié: (2026)
par: Nicolae, Radu, et autres
Publié: (2026)
KV Cache Compression for Inference Efficiency in LLMs: A Review
par: Liu, Yanyu, et autres
Publié: (2025)
par: Liu, Yanyu, et autres
Publié: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
par: Xu, Chuhao, et autres
Publié: (2025)
par: Xu, Chuhao, et autres
Publié: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
par: Lin, Haoran, et autres
Publié: (2025)
par: Lin, Haoran, et autres
Publié: (2025)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
par: Lee, Sunjung, et autres
Publié: (2026)
par: Lee, Sunjung, et autres
Publié: (2026)
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
par: Lee, Wonbeom, et autres
Publié: (2024)
par: Lee, Wonbeom, et autres
Publié: (2024)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
par: Zhong, Shuzhang, et autres
Publié: (2025)
par: Zhong, Shuzhang, et autres
Publié: (2025)
Parallax: Efficient LLM Inference Service over Decentralized Environment
par: Tong, Chris, et autres
Publié: (2025)
par: Tong, Chris, et autres
Publié: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
par: He, Yiyuan, et autres
Publié: (2024)
par: He, Yiyuan, et autres
Publié: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
par: He, Wenhao, et autres
Publié: (2026)
par: He, Wenhao, et autres
Publié: (2026)
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
par: Sun, Minqiu, et autres
Publié: (2026)
par: Sun, Minqiu, et autres
Publié: (2026)
CacheFL: Privacy-Preserving and Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
par: Yi, Mengjun, et autres
Publié: (2025)
par: Yi, Mengjun, et autres
Publié: (2025)
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
par: Ruan, Chaoyi, et autres
Publié: (2025)
par: Ruan, Chaoyi, et autres
Publié: (2025)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
par: Zhao, Alan, et autres
Publié: (2026)
par: Zhao, Alan, et autres
Publié: (2026)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
par: Zhang, Mingjin, et autres
Publié: (2024)
par: Zhang, Mingjin, et autres
Publié: (2024)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
par: Zhang, Songge, et autres
Publié: (2026)
par: Zhang, Songge, et autres
Publié: (2026)
Asynchronous Checkpoint for Eventually Consistent Databases
par: Ravishankar, Raaghav, et autres
Publié: (2025)
par: Ravishankar, Raaghav, et autres
Publié: (2025)
Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
par: Zhu, Yue, et autres
Publié: (2025)
par: Zhu, Yue, et autres
Publié: (2025)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
par: Kong, Jie, et autres
Publié: (2026)
par: Kong, Jie, et autres
Publié: (2026)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
par: Sandholm, Thomas, et autres
Publié: (2025)
par: Sandholm, Thomas, et autres
Publié: (2025)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
par: Yu, Shibo, et autres
Publié: (2025)
par: Yu, Shibo, et autres
Publié: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
par: Jeong, Bodon, et autres
Publié: (2026)
par: Jeong, Bodon, et autres
Publié: (2026)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
par: Gan, Zhenghao, et autres
Publié: (2026)
par: Gan, Zhenghao, et autres
Publié: (2026)
InstCache: A Predictive Cache for LLM Serving
par: Zou, Longwei, et autres
Publié: (2024)
par: Zou, Longwei, et autres
Publié: (2024)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
par: Stoyanov, Radostin, et autres
Publié: (2025)
par: Stoyanov, Radostin, et autres
Publié: (2025)
Scrutinizing Variables for Checkpoint Using Automatic Differentiation
par: Huang, Xin, et autres
Publié: (2026)
par: Huang, Xin, et autres
Publié: (2026)
Optimal Checkpoint Interval with Availability as an Objective Function
par: Saxena, Nirmal Raj, et autres
Publié: (2024)
par: Saxena, Nirmal Raj, et autres
Publié: (2024)
Checkpoint and Restart: An Energy Consumption Characterization in Clusters
par: Moran, Marina, et autres
Publié: (2024)
par: Moran, Marina, et autres
Publié: (2024)
Documents similaires
-
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
par: Cho, Jaehong, et autres
Publié: (2024) -
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
par: Cho, Jaehong, et autres
Publié: (2026) -
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
par: Yoon, Dongha, et autres
Publié: (2025) -
Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
par: Hwang, Jinwoo, et autres
Publié: (2025) -
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
par: Nian, Sean, et autres
Publié: (2026)