Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Zixu, Wei, Jianyu, Wang, Tuowei, Huang, Minxing, Jiang, Huiqiang, Jiang, Shiqi, Cao, Ting, Ren, Ju |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
by: Wang, Tuowei, et al.
Published: (2025)
by: Wang, Tuowei, et al.
Published: (2025)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
WindVE: Collaborative CPU-NPU Vector Embedding
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
UMDAM: A Unified Data Layout and DRAM Address Mapping for Heterogenous NPU-PIM
by: Huang, Hai
Published: (2025)
by: Huang, Hai
Published: (2025)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
by: Tummalapalli, Pranay, et al.
Published: (2026)
by: Tummalapalli, Pranay, et al.
Published: (2026)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
by: Dai, Yuntao, et al.
Published: (2026)
by: Dai, Yuntao, et al.
Published: (2026)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
by: Chen, Kefu, et al.
Published: (2026)
by: Chen, Kefu, et al.
Published: (2026)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
by: Xue, Chunyu, et al.
Published: (2026)
by: Xue, Chunyu, et al.
Published: (2026)
T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
by: Wei, Jianyu, et al.
Published: (2024)
by: Wei, Jianyu, et al.
Published: (2024)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
by: Lai, Ruiqi, et al.
Published: (2025)
by: Lai, Ruiqi, et al.
Published: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
by: Li, Xiangyu, et al.
Published: (2025)
by: Li, Xiangyu, et al.
Published: (2025)
e112: A Context-Aware Mobile Emergency Communication Platform Leveraging Smartphone Sensing and Cloud Services
by: Ioannidou, Katerina, et al.
Published: (2026)
by: Ioannidou, Katerina, et al.
Published: (2026)
Mobile Edge Computing
by: Ahmed, Sohaib, et al.
Published: (2024)
by: Ahmed, Sohaib, et al.
Published: (2024)
A Joint Time and Energy-Efficient Federated Learning-based Computation Offloading Method for Mobile Edge Computing
by: Mukherjee, Anwesha, et al.
Published: (2024)
by: Mukherjee, Anwesha, et al.
Published: (2024)
Resource-efficient Parallel Split Learning in Heterogeneous Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
Energy-Efficient Real-Time Job Mapping and Resource Management in Mobile-Edge Computing
by: Gao, Chuanchao, et al.
Published: (2025)
by: Gao, Chuanchao, et al.
Published: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
by: Wei, Jinhui, et al.
Published: (2025)
by: Wei, Jinhui, et al.
Published: (2025)
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
by: Scheinert, Dominik, et al.
Published: (2026)
by: Scheinert, Dominik, et al.
Published: (2026)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
Communication-and-Computation Efficient Split Federated Learning: Gradient Aggregation and Resource Management
by: Liang, Yipeng, et al.
Published: (2025)
by: Liang, Yipeng, et al.
Published: (2025)
LLM-Enhanced Deep Reinforcement Learning for Task Offloading in Collaborative Edge Computing
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
by: Zhao, Xinkui, et al.
Published: (2025)
by: Zhao, Xinkui, et al.
Published: (2025)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
by: Chen, Jiu, et al.
Published: (2026)
by: Chen, Jiu, et al.
Published: (2026)
FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
by: Ma, Jian, et al.
Published: (2025)
by: Ma, Jian, et al.
Published: (2025)
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
by: Zhao, Juntao, et al.
Published: (2025)
by: Zhao, Juntao, et al.
Published: (2025)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
by: Zhang, Hongbin, et al.
Published: (2026)
by: Zhang, Hongbin, et al.
Published: (2026)
Barycentric Coded Distributed Computing with Flexible Recovery Threshold for Collaborative Mobile Edge Computing
by: Qiu, Houming, et al.
Published: (2025)
by: Qiu, Houming, et al.
Published: (2025)
From Patchwork to Network: A Comprehensive Framework for Demand Analysis and Fleet Optimization of Urban Air Mobility
by: Jiang, Xuan, et al.
Published: (2025)
by: Jiang, Xuan, et al.
Published: (2025)
Blockchain-Enabled Dynamic Spectrum Sharing for Satellite and Terrestrial Communication Networks
by: Wang, Zixin, et al.
Published: (2024)
by: Wang, Zixin, et al.
Published: (2024)
Computational Power of Mobile Robots in Synchronous Environment: Discrete Version
by: Sharma, Avisek, et al.
Published: (2024)
by: Sharma, Avisek, et al.
Published: (2024)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
by: Zhao, Alan, et al.
Published: (2026)
by: Zhao, Alan, et al.
Published: (2026)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
CCRSat: A Collaborative Computation Reuse Framework for Satellite Edge Computing Networks
by: Zhang, Ye, et al.
Published: (2025)
by: Zhang, Ye, et al.
Published: (2025)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
by: Peng, You, et al.
Published: (2026)
by: Peng, You, et al.
Published: (2026)
Mobile Cloud Computing in Healthcare Using Dynamic Cloudlets for Energy-Aware Consumption
by: Muniswamaiah, Manoj, et al.
Published: (2019)
by: Muniswamaiah, Manoj, et al.
Published: (2019)
EcoLife: Carbon-Aware Serverless Function Scheduling for Sustainable Computing
by: Jiang, Yankai, et al.
Published: (2024)
by: Jiang, Yankai, et al.
Published: (2024)
BEFL: Balancing Energy Consumption in Federated Learning for Mobile Edge IoT
by: Ju, Zehao, et al.
Published: (2024)
by: Ju, Zehao, et al.
Published: (2024)
Exploring Uncore Frequency Scaling for Heterogeneous Computing
by: Zheng, Zhong, et al.
Published: (2025)
by: Zheng, Zhong, et al.
Published: (2025)
Similar Items
-
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
by: Wang, Tuowei, et al.
Published: (2025) -
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
by: Wang, Tuowei, et al.
Published: (2024) -
WindVE: Collaborative CPU-NPU Vector Embedding
by: Huang, Jinqi, et al.
Published: (2025) -
UMDAM: A Unified Data Layout and DRAM Address Mapping for Heterogenous NPU-PIM
by: Huang, Hai
Published: (2025) -
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
by: Tummalapalli, Pranay, et al.
Published: (2026)