Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Zongpu, Dash, Pranab, Hu, Y. Charlie, Xu, Qiang, Li, Jian, Guan, Haibing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RTP-LLM: High-Performance Alibaba LLM Inference Engine
di: Tan, Boyu, et al.
Pubblicazione: (2026)
di: Tan, Boyu, et al.
Pubblicazione: (2026)
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
di: Liu, Jinshu, et al.
Pubblicazione: (2024)
di: Liu, Jinshu, et al.
Pubblicazione: (2024)
Energy-Efficient Computation with DVFS using Deep Reinforcement Learning for Multi-Task Systems in Edge Computing
di: Li, Xinyi, et al.
Pubblicazione: (2024)
di: Li, Xinyi, et al.
Pubblicazione: (2024)
LLM as a System Service on Mobile Devices
di: Yin, Wangsong, et al.
Pubblicazione: (2024)
di: Yin, Wangsong, et al.
Pubblicazione: (2024)
AIOS: LLM Agent Operating System
di: Mei, Kai, et al.
Pubblicazione: (2024)
di: Mei, Kai, et al.
Pubblicazione: (2024)
VeriLocc: End-to-End Cross-Architecture Register Allocation via LLM
di: Jin, Lesheng, et al.
Pubblicazione: (2025)
di: Jin, Lesheng, et al.
Pubblicazione: (2025)
SSV: Sparse Speculative Verification for Efficient LLM Inference
di: Wang, Zhibin, et al.
Pubblicazione: (2026)
di: Wang, Zhibin, et al.
Pubblicazione: (2026)
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
di: Du, Hongchao, et al.
Pubblicazione: (2025)
di: Du, Hongchao, et al.
Pubblicazione: (2025)
GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
di: Zhang, Keyao, et al.
Pubblicazione: (2025)
di: Zhang, Keyao, et al.
Pubblicazione: (2025)
Potential of WebAssembly for Embedded Systems
di: Wallentowitz, Stefan, et al.
Pubblicazione: (2024)
di: Wallentowitz, Stefan, et al.
Pubblicazione: (2024)
OBASE: Object-Based Address-Space Engineering to Improve Memory Tiering
di: Banakar, Vinay, et al.
Pubblicazione: (2026)
di: Banakar, Vinay, et al.
Pubblicazione: (2026)
Getting a Handle on Unmanaged Memory
di: Wanninger, Nick, et al.
Pubblicazione: (2024)
di: Wanninger, Nick, et al.
Pubblicazione: (2024)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
di: Qiu, Shi, et al.
Pubblicazione: (2026)
di: Qiu, Shi, et al.
Pubblicazione: (2026)
Assessing FIFO and Round Robin Scheduling:Effects on Data Pipeline Performance and Energy Usage
di: Choudhury, Malobika Roy, et al.
Pubblicazione: (2024)
di: Choudhury, Malobika Roy, et al.
Pubblicazione: (2024)
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
di: Cui, Weihao, et al.
Pubblicazione: (2025)
di: Cui, Weihao, et al.
Pubblicazione: (2025)
Principled Performance Tunability in Operating System Kernels
di: Chen, Zhongjie, et al.
Pubblicazione: (2025)
di: Chen, Zhongjie, et al.
Pubblicazione: (2025)
Scaling Inter-procedural Dataflow Analysis on the Cloud
di: Sun, Zewen, et al.
Pubblicazione: (2024)
di: Sun, Zewen, et al.
Pubblicazione: (2024)
Tidying Up the Address Space
di: Banakar, Vinay, et al.
Pubblicazione: (2025)
di: Banakar, Vinay, et al.
Pubblicazione: (2025)
Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate
di: Liu, Fangyue, et al.
Pubblicazione: (2026)
di: Liu, Fangyue, et al.
Pubblicazione: (2026)
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
di: Kang, Hao, et al.
Pubblicazione: (2026)
di: Kang, Hao, et al.
Pubblicazione: (2026)
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
di: Huang, Zhengxiang, et al.
Pubblicazione: (2025)
di: Huang, Zhengxiang, et al.
Pubblicazione: (2025)
WebAssembly on Resource-Constrained IoT Devices: Performance, Efficiency, and Portability
di: Has, Mislav, et al.
Pubblicazione: (2025)
di: Has, Mislav, et al.
Pubblicazione: (2025)
Quine: Realizing LLM Agents as Native POSIX Processes
di: Ke, Hao
Pubblicazione: (2026)
di: Ke, Hao
Pubblicazione: (2026)
Cerebrum (AIOS SDK): A Platform for Agent Development, Deployment, Distribution, and Discovery
di: Rama, Balaji, et al.
Pubblicazione: (2025)
di: Rama, Balaji, et al.
Pubblicazione: (2025)
Scalable and Accurate Application-Level Crash-Consistency Testing via Representative Testing
di: Gu, Yile, et al.
Pubblicazione: (2025)
di: Gu, Yile, et al.
Pubblicazione: (2025)
Horizon-LM: A RAM-Centric Architecture for LLM Training
di: Yuan, Zhengqing, et al.
Pubblicazione: (2026)
di: Yuan, Zhengqing, et al.
Pubblicazione: (2026)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
di: Chen, Yukang, et al.
Pubblicazione: (2025)
di: Chen, Yukang, et al.
Pubblicazione: (2025)
Revitalising the Single Batch Environment: A 'Quest' to Achieve Fairness and Efficiency
di: Manna, Supriya, et al.
Pubblicazione: (2023)
di: Manna, Supriya, et al.
Pubblicazione: (2023)
Semantic Scheduling for LLM Inference
di: Hua, Wenyue, et al.
Pubblicazione: (2025)
di: Hua, Wenyue, et al.
Pubblicazione: (2025)
Decoupling Vector Data and Index Storage for Space Efficiency
di: Ren, Yuanming, et al.
Pubblicazione: (2026)
di: Ren, Yuanming, et al.
Pubblicazione: (2026)
Towards Agentic OS: An LLM Agent Framework for Linux Schedulers
di: Zheng, Yusheng, et al.
Pubblicazione: (2025)
di: Zheng, Yusheng, et al.
Pubblicazione: (2025)
Vmem: A Lightweight Hot-Upgradable Memory Management for In-production Cloud Environment
di: Zheng, Hao, et al.
Pubblicazione: (2025)
di: Zheng, Hao, et al.
Pubblicazione: (2025)
Compiling Away the Overhead of Race Detection
di: Paznikov, Alexey, et al.
Pubblicazione: (2025)
di: Paznikov, Alexey, et al.
Pubblicazione: (2025)
Sockeye: a language for analyzing hardware documentation
di: Fiedler, Ben, et al.
Pubblicazione: (2025)
di: Fiedler, Ben, et al.
Pubblicazione: (2025)
Futureproof Static Memory Planning
di: Lamprakos, Christos, et al.
Pubblicazione: (2025)
di: Lamprakos, Christos, et al.
Pubblicazione: (2025)
vNV-Heap: An Ownership-Based Virtually Non-Volatile Heap for Embedded Systems
di: Gerber, Markus Elias, et al.
Pubblicazione: (2025)
di: Gerber, Markus Elias, et al.
Pubblicazione: (2025)
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
di: Luo, Shutian, et al.
Pubblicazione: (2026)
di: Luo, Shutian, et al.
Pubblicazione: (2026)
ASC-Hook: fast and transparent system call hook for Arm
di: Shen, Yang, et al.
Pubblicazione: (2024)
di: Shen, Yang, et al.
Pubblicazione: (2024)
Clove: Object-Level CXL Memory Management in Managed Runtimes
di: Son, Sam, et al.
Pubblicazione: (2026)
di: Son, Sam, et al.
Pubblicazione: (2026)
Safe and usable kernel extensions with Rex
di: Jia, Jinghao, et al.
Pubblicazione: (2025)
di: Jia, Jinghao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RTP-LLM: High-Performance Alibaba LLM Inference Engine
di: Tan, Boyu, et al.
Pubblicazione: (2026) -
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
di: Liu, Jinshu, et al.
Pubblicazione: (2024) -
Energy-Efficient Computation with DVFS using Deep Reinforcement Learning for Multi-Task Systems in Edge Computing
di: Li, Xinyi, et al.
Pubblicazione: (2024) -
LLM as a System Service on Mobile Devices
di: Yin, Wangsong, et al.
Pubblicazione: (2024) -
AIOS: LLM Agent Operating System
di: Mei, Kai, et al.
Pubblicazione: (2024)