PIPO: Pipelined Offloading for Efficient Inference on Consumer Devices
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yangyijian, Li, Jun, Li, Wu-Jun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
di: Han, Xueyuan, et al.
Pubblicazione: (2024)
di: Han, Xueyuan, et al.
Pubblicazione: (2024)
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
di: Wan, Xinyi, et al.
Pubblicazione: (2025)
di: Wan, Xinyi, et al.
Pubblicazione: (2025)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
di: Moothedath, Vishnu Narayanan, et al.
Pubblicazione: (2025)
di: Moothedath, Vishnu Narayanan, et al.
Pubblicazione: (2025)
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
di: Wang, Kun, et al.
Pubblicazione: (2024)
di: Wang, Kun, et al.
Pubblicazione: (2024)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
di: Jiang, Xuanlin, et al.
Pubblicazione: (2024)
di: Jiang, Xuanlin, et al.
Pubblicazione: (2024)
DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
di: Peng, Xian, et al.
Pubblicazione: (2025)
di: Peng, Xian, et al.
Pubblicazione: (2025)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
di: Dege, Pengcuo, et al.
Pubblicazione: (2025)
di: Dege, Pengcuo, et al.
Pubblicazione: (2025)
AgentStop: Terminating Local AI Agents Early to Save Energy in Consumer Devices
di: Pham, Dzung, et al.
Pubblicazione: (2026)
di: Pham, Dzung, et al.
Pubblicazione: (2026)
Efficient Training on Multiple Consumer GPUs with RoundPipe
di: Luo, Yibin, et al.
Pubblicazione: (2026)
di: Luo, Yibin, et al.
Pubblicazione: (2026)
Backpropagation-Free Multi-modal On-Device Model Adaptation via Cloud-Device Collaboration
di: Ji, Wei, et al.
Pubblicazione: (2024)
di: Ji, Wei, et al.
Pubblicazione: (2024)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
di: Shen, Zheyu, et al.
Pubblicazione: (2025)
di: Shen, Zheyu, et al.
Pubblicazione: (2025)
Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
di: Li, Yang, et al.
Pubblicazione: (2025)
di: Li, Yang, et al.
Pubblicazione: (2025)
Adaptive and Resource-efficient Agentic AI Systems for Mobile and Embedded Devices: A Survey
di: Liu, Sicong, et al.
Pubblicazione: (2025)
di: Liu, Sicong, et al.
Pubblicazione: (2025)
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
di: Li, Yan, et al.
Pubblicazione: (2025)
di: Li, Yan, et al.
Pubblicazione: (2025)
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
di: Yuan, Ziqi, et al.
Pubblicazione: (2025)
di: Yuan, Ziqi, et al.
Pubblicazione: (2025)
Efficient Federated Finetuning of Tiny Transformers with Resource-Constrained Devices
di: Pfeiffer, Kilian, et al.
Pubblicazione: (2024)
di: Pfeiffer, Kilian, et al.
Pubblicazione: (2024)
Learning Like Humans: Resource-Efficient Federated Fine-Tuning through Cognitive Developmental Stages
di: Wu, Yebo, et al.
Pubblicazione: (2025)
di: Wu, Yebo, et al.
Pubblicazione: (2025)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
di: Wu, Houming, et al.
Pubblicazione: (2024)
di: Wu, Houming, et al.
Pubblicazione: (2024)
PNCS:Power-Norm Cosine Similarity for Diverse Client Selection in Federated Learning
di: Li, Liangyan, et al.
Pubblicazione: (2025)
di: Li, Liangyan, et al.
Pubblicazione: (2025)
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
di: Dash, Sajal, et al.
Pubblicazione: (2026)
di: Dash, Sajal, et al.
Pubblicazione: (2026)
Zero Bubble Pipeline Parallelism
di: Qi, Penghui, et al.
Pubblicazione: (2023)
di: Qi, Penghui, et al.
Pubblicazione: (2023)
TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training
di: Wu, Houming, et al.
Pubblicazione: (2025)
di: Wu, Houming, et al.
Pubblicazione: (2025)
When Foresight Pruning Meets Zeroth-Order Optimization: Efficient Federated Learning for Low-Memory Devices
di: Zhang, Pengyu, et al.
Pubblicazione: (2024)
di: Zhang, Pengyu, et al.
Pubblicazione: (2024)
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
di: Xie, Jianhang, et al.
Pubblicazione: (2025)
di: Xie, Jianhang, et al.
Pubblicazione: (2025)
SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment
di: Zhu, Wenqiao, et al.
Pubblicazione: (2025)
di: Zhu, Wenqiao, et al.
Pubblicazione: (2025)
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
di: Gu, Yile, et al.
Pubblicazione: (2025)
di: Gu, Yile, et al.
Pubblicazione: (2025)
From Centralized to Decentralized Federated Learning: Theoretical Insights, Privacy Preservation, and Robustness Challenges
di: Li, Qiongxiu, et al.
Pubblicazione: (2025)
di: Li, Qiongxiu, et al.
Pubblicazione: (2025)
Local-Cloud Inference Offloading for LLMs in Multi-Modal, Multi-Task, Multi-Dialogue Settings
di: Yuan, Liangqi, et al.
Pubblicazione: (2025)
di: Yuan, Liangqi, et al.
Pubblicazione: (2025)
Learn How to Query from Unlabeled Data Streams in Federated Learning
di: Sun, Yuchang, et al.
Pubblicazione: (2024)
di: Sun, Yuchang, et al.
Pubblicazione: (2024)
Synera: Synergistic LLM Serving across Device and Cloud at Scale
di: Wang, Genglin, et al.
Pubblicazione: (2025)
di: Wang, Genglin, et al.
Pubblicazione: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
di: Ye, Zihao, et al.
Pubblicazione: (2025)
di: Ye, Zihao, et al.
Pubblicazione: (2025)
A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing
di: Seal, Sudip K., et al.
Pubblicazione: (2025)
di: Seal, Sudip K., et al.
Pubblicazione: (2025)
Distributed Graph Neural Network Inference With Just-In-Time Compilation For Industry-Scale Graphs
di: Wu, Xiabao, et al.
Pubblicazione: (2025)
di: Wu, Xiabao, et al.
Pubblicazione: (2025)
DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
di: Zhu, Zeyu, et al.
Pubblicazione: (2026)
di: Zhu, Zeyu, et al.
Pubblicazione: (2026)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
di: Yu, Hanfei, et al.
Pubblicazione: (2025)
di: Yu, Hanfei, et al.
Pubblicazione: (2025)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
di: Jeon, Byungsoo, et al.
Pubblicazione: (2024)
di: Jeon, Byungsoo, et al.
Pubblicazione: (2024)
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
di: Yu, Donglin
Pubblicazione: (2026)
di: Yu, Donglin
Pubblicazione: (2026)
A Survey on Inference Optimization Techniques for Mixture of Experts Models
di: Liu, Jiacheng, et al.
Pubblicazione: (2024)
di: Liu, Jiacheng, et al.
Pubblicazione: (2024)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
di: Wang, Taiyi, et al.
Pubblicazione: (2024)
di: Wang, Taiyi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
di: Han, Xueyuan, et al.
Pubblicazione: (2024) -
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
di: Wan, Xinyi, et al.
Pubblicazione: (2025) -
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
di: Moothedath, Vishnu Narayanan, et al.
Pubblicazione: (2025) -
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
di: Wang, Kun, et al.
Pubblicazione: (2024) -
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
di: Jiang, Xuanlin, et al.
Pubblicazione: (2024)