MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Zhengxiang, Niu, Chaoyue, Wang, Zhaode, Xue, Jiarui, Zhang, Hanming, Wang, Yugang, Xin, Zewei, Jiang, Xiaotang, Lv, Chengfei, Wu, Fan, Chen, Guihai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices
by: Wang, Zhaode, et al.
Published: (2025)
by: Wang, Zhaode, et al.
Published: (2025)
2DIO: A Cache-Accurate Storage Microbenchmark
by: Wang, Yirong, et al.
Published: (2026)
by: Wang, Yirong, et al.
Published: (2026)
A System-Level Dynamic Binary Translator using Automatically-Learned Translation Rules
by: Jiang, Jinhu, et al.
Published: (2024)
by: Jiang, Jinhu, et al.
Published: (2024)
Delegation with Trust<T>: A Scalable, Type- and Memory-Safe Alternative to Locks
by: Ahmad, Noaman, et al.
Published: (2024)
by: Ahmad, Noaman, et al.
Published: (2024)
Columbo: Low Level End-to-End System Traces through Modular Full-System Simulation
by: Görgen, Jakob, et al.
Published: (2024)
by: Görgen, Jakob, et al.
Published: (2024)
Inspection of I/O Operations from System Call Traces using Directly-Follows-Graph
by: Sankaran, Aravind, et al.
Published: (2024)
by: Sankaran, Aravind, et al.
Published: (2024)
CORD: Co-design of Resource Allocation and Deadline Decomposition with Generative Profiling
by: Gifford, Robert, et al.
Published: (2025)
by: Gifford, Robert, et al.
Published: (2025)
Optimizing System Memory Bandwidth with Micron CXL Memory Expansion Modules on Intel Xeon 6 Processors
by: Sehgal, Rohit, et al.
Published: (2024)
by: Sehgal, Rohit, et al.
Published: (2024)
Performance Characterization of AutoNUMA Memory Tiering on Graph Analytics
by: Moura, Diego, et al.
Published: (2022)
by: Moura, Diego, et al.
Published: (2022)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counters Tracking
by: Mazzola, Sergio, et al.
Published: (2024)
by: Mazzola, Sergio, et al.
Published: (2024)
Energy-Aware CPU Orchestration in O-RAN: A dApp-Driven Lightweight Approach
by: Crespo, Francisco, et al.
Published: (2025)
by: Crespo, Francisco, et al.
Published: (2025)
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
by: Liargkovas, Georgios, et al.
Published: (2026)
by: Liargkovas, Georgios, et al.
Published: (2026)
CXLMemSim: A pure software simulated CXL.mem for performance characterization
by: Yang, Yiwei, et al.
Published: (2023)
by: Yang, Yiwei, et al.
Published: (2023)
Tidying Up the Address Space
by: Banakar, Vinay, et al.
Published: (2025)
by: Banakar, Vinay, et al.
Published: (2025)
Putting the Context back into Memory
by: Roberts, David A.
Published: (2025)
by: Roberts, David A.
Published: (2025)
A Limits Study of Memory-side Tiering Telemetry
by: Petrucci, Vinicius, et al.
Published: (2025)
by: Petrucci, Vinicius, et al.
Published: (2025)
CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural Assumptions (Extended Version)
by: Lindsay, Nick, et al.
Published: (2026)
by: Lindsay, Nick, et al.
Published: (2026)
Random Adaptive Cache Placement Policy
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
Sensifi: A Wireless Sensing System for Ultra-High-Rate Applications
by: Li, Chia-Chi, et al.
Published: (2020)
by: Li, Chia-Chi, et al.
Published: (2020)
Characterizing Physical Memory Fragmentation
by: Mansi, Mark, et al.
Published: (2024)
by: Mansi, Mark, et al.
Published: (2024)
Mitigating GIL Bottlenecks in Edge AI Systems
by: Mandal, Mridankan, et al.
Published: (2026)
by: Mandal, Mridankan, et al.
Published: (2026)
ASC-Hook: fast and transparent system call hook for Arm
by: Shen, Yang, et al.
Published: (2024)
by: Shen, Yang, et al.
Published: (2024)
RAID Organizations for Improved Reliability and Performance: A Not Entirely Unbiased Tutorial (1st revision)
by: Thomasian, Alexander
Published: (2024)
by: Thomasian, Alexander
Published: (2024)
SwitchFS: Asynchronous Metadata Updates for Distributed Filesystems with In-Network Coordination
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
CPU-Limits kill Performance: Time to rethink Resource Control
by: Shetty, Chirag, et al.
Published: (2025)
by: Shetty, Chirag, et al.
Published: (2025)
A TRRIP Down Memory Lane: Temperature-Based Re-Reference Interval Prediction For Instruction Caching
by: Kao, Henry, et al.
Published: (2025)
by: Kao, Henry, et al.
Published: (2025)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Demystifying Serverless Costs on Public Platforms: Bridging Billing, Architecture, and OS Scheduling
by: Lin, Changyuan, et al.
Published: (2025)
by: Lin, Changyuan, et al.
Published: (2025)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
by: Tofigh, Mani, et al.
Published: (2025)
by: Tofigh, Mani, et al.
Published: (2025)
GPUVM: GPU-driven Unified Virtual Memory
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
The Impact of Private Equity and Venture Capital Funds on post-IPO Operational and Financial Performance in Brazilian invested companies
by: Bianca Piloto Sincerre
Published: (2019)
by: Bianca Piloto Sincerre
Published: (2019)
Understanding and Enhancing Linux Kernel-based Packet Switching on WiFi Access Points
by: Zhang, Shiqi, et al.
Published: (2024)
by: Zhang, Shiqi, et al.
Published: (2024)
The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
by: Xu, Yuhang, et al.
Published: (2025)
by: Xu, Yuhang, et al.
Published: (2025)
A Case for CATS: A Conductor-driven Asymmetric Transport Scheme for Semantic Prioritization
by: Rizvi, Syed Muhammad Aqdas
Published: (2026)
by: Rizvi, Syed Muhammad Aqdas
Published: (2026)
Serverless Cold Starts and Where to Find Them
by: Joosen, Artjom, et al.
Published: (2024)
by: Joosen, Artjom, et al.
Published: (2024)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
Optimizing Airline Customer Service: A KPI-Driven Approach for Chief Customer Services Officers
by: MoghadasNian, SeyyedAbdolHojjat, et al.
Published: (2024)
by: MoghadasNian, SeyyedAbdolHojjat, et al.
Published: (2024)
Similar Items
-
MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices
by: Wang, Zhaode, et al.
Published: (2025) -
2DIO: A Cache-Accurate Storage Microbenchmark
by: Wang, Yirong, et al.
Published: (2026) -
A System-Level Dynamic Binary Translator using Automatically-Learned Translation Rules
by: Jiang, Jinhu, et al.
Published: (2024) -
Delegation with Trust<T>: A Scalable, Type- and Memory-Safe Alternative to Locks
by: Ahmad, Noaman, et al.
Published: (2024) -
Columbo: Low Level End-to-End System Traces through Modular Full-System Simulation
by: Görgen, Jakob, et al.
Published: (2024)