InstMeter: An Instruction-Level Method to Predict Energy and Latency of DL Model Inference on MCUs
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Hao, Wang, Qing, Zuniga, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
by: Wang, Haoxin, et al.
Published: (2025)
by: Wang, Haoxin, et al.
Published: (2025)
Low-Energy On-Device Personalization for MCUs
by: Huang, Yushan, et al.
Published: (2024)
by: Huang, Yushan, et al.
Published: (2024)
UnIT: Scalable Unstructured Inference-Time Pruning for MAC-efficient Neural Inference on MCUs
by: Neth, Ashe, et al.
Published: (2025)
by: Neth, Ashe, et al.
Published: (2025)
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
by: Zheng, Size, et al.
Published: (2024)
by: Zheng, Size, et al.
Published: (2024)
ML Inference Scheduling with Predictable Latency
by: Zhao, Haidong, et al.
Published: (2025)
by: Zhao, Haidong, et al.
Published: (2025)
MONAS: Efficient Zero-Shot Neural Architecture Search for MCUs
by: Qiao, Ye, et al.
Published: (2024)
by: Qiao, Ye, et al.
Published: (2024)
MicroNAS: Zero-Shot Neural Architecture Search for MCUs
by: Qiao, Ye, et al.
Published: (2024)
by: Qiao, Ye, et al.
Published: (2024)
MambaLite-Micro: Memory-Optimized Mamba Inference on MCUs
by: Xu, Hongjun, et al.
Published: (2025)
by: Xu, Hongjun, et al.
Published: (2025)
Divide-Conquer Transformer Learning for Predicting Electric Vehicle Charging Events Using Smart Meter Data
by: Ke, Fucai, et al.
Published: (2024)
by: Ke, Fucai, et al.
Published: (2024)
Latenrgy: Model Agnostic Latency and Energy Consumption Prediction for Binary Classifiers
by: Pittman, Jason M.
Published: (2024)
by: Pittman, Jason M.
Published: (2024)
GraphPPD: Posterior Predictive Modelling for Graph-Level Inference
by: Pal, Soumyasundar, et al.
Published: (2025)
by: Pal, Soumyasundar, et al.
Published: (2025)
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
by: Jung, Victor J. B., et al.
Published: (2024)
by: Jung, Victor J. B., et al.
Published: (2024)
ScaleDL: Towards Scalable and Efficient Runtime Prediction for Distributed Deep Learning Workloads
by: Wang, Xiaokai, et al.
Published: (2025)
by: Wang, Xiaokai, et al.
Published: (2025)
SmartMeterFM: Unifying Smart Meter Data Generative Tasks Using Flow Matching Models
by: Lin, Nan, et al.
Published: (2026)
by: Lin, Nan, et al.
Published: (2026)
Benchmarking Energy and Latency in TinyML: A Novel Method for Resource-Constrained AI
by: Bartoli, Pietro, et al.
Published: (2025)
by: Bartoli, Pietro, et al.
Published: (2025)
MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees
by: Zhang, Ryan, et al.
Published: (2024)
by: Zhang, Ryan, et al.
Published: (2024)
A Study on Inference Latency for Vision Transformers on Mobile Devices
by: Li, Zhuojin, et al.
Published: (2025)
by: Li, Zhuojin, et al.
Published: (2025)
Fast and Compact Tsetlin Machine Inference on CPUs Using Instruction-Level Optimization
by: Zeng, Yefan, et al.
Published: (2025)
by: Zeng, Yefan, et al.
Published: (2025)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
by: Sui, Yifan, et al.
Published: (2025)
by: Sui, Yifan, et al.
Published: (2025)
RepDL: Bit-level Reproducible Deep Learning Training and Inference
by: Xie, Peichen, et al.
Published: (2025)
by: Xie, Peichen, et al.
Published: (2025)
Stateful Inference for Low-Latency Multi-Agent Tool Calling
by: Norgren, Victor
Published: (2026)
by: Norgren, Victor
Published: (2026)
Inference Energy and Latency in AI-Mediated Education: A Learning-per-Watt Analysis of Edge and Cloud Models
by: Khemani, Kushal
Published: (2026)
by: Khemani, Kushal
Published: (2026)
A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
A Data Mining-Based Dynamical Anomaly Detection Method for Integrating with an Advance Metering System
by: Maitra, Sarit
Published: (2024)
by: Maitra, Sarit
Published: (2024)
DVFS-Aware DNN Inference on GPUs: Latency Modeling and Performance Analysis
by: Han, Yunchu, et al.
Published: (2025)
by: Han, Yunchu, et al.
Published: (2025)
Fully Autonomous Z-Score-Based TinyML Anomaly Detection on Resource-Constrained MCUs Using Power Side-Channel Data
by: Albaiz, Abdulrahman, et al.
Published: (2026)
by: Albaiz, Abdulrahman, et al.
Published: (2026)
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
by: Zhao, Minjun, et al.
Published: (2023)
by: Zhao, Minjun, et al.
Published: (2023)
MCU-MixQ: A HW/SW Co-optimized Mixed-precision Neural Network Design Framework for MCUs
by: Gong, Junfeng, et al.
Published: (2024)
by: Gong, Junfeng, et al.
Published: (2024)
The Spectral Geometry of Thought: Phase Transitions, Instruction Reversal, Token-Level Dynamics, and Perfect Correctness Prediction in How Transformers Reason
by: Liu, Yi
Published: (2026)
by: Liu, Yi
Published: (2026)
Low Latency Transformer Inference on FPGAs for Physics Applications with hls4ml
by: Jiang, Zhixing, et al.
Published: (2024)
by: Jiang, Zhixing, et al.
Published: (2024)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
by: Husom, Erik Johannes, et al.
Published: (2025)
by: Husom, Erik Johannes, et al.
Published: (2025)
Benchmarking of EEG Analysis Techniques for Parkinson's Disease Diagnosis: A Comparison between Traditional ML Methods and Foundation DL Methods
by: Avola, Danilo, et al.
Published: (2025)
by: Avola, Danilo, et al.
Published: (2025)
Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models
by: Delavande, Julien, et al.
Published: (2025)
by: Delavande, Julien, et al.
Published: (2025)
Towards Secure and Scalable Energy Theft Detection: A Federated Learning Approach for Resource-Constrained Smart Meters
by: Labate, Diego, et al.
Published: (2026)
by: Labate, Diego, et al.
Published: (2026)
Green AI: A Preliminary Empirical Study on Energy Consumption in DL Models Across Different Runtime Infrastructures
by: Alizadeh, Negar, et al.
Published: (2024)
by: Alizadeh, Negar, et al.
Published: (2024)
EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
by: Song, Zhiye, et al.
Published: (2026)
by: Song, Zhiye, et al.
Published: (2026)
Ti-iLSTM: A TinyDL Approach for Logic-Level Anomaly Detection in Industrial Water Treatment Systems
by: Joshi, Mandar, et al.
Published: (2026)
by: Joshi, Mandar, et al.
Published: (2026)
Outcome-Aware Tool Selection for Semantic Routers: Latency-Constrained Learning Without LLM Inference
by: Chen, Huamin, et al.
Published: (2026)
by: Chen, Huamin, et al.
Published: (2026)
Similar Items
-
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
by: Wang, Haoxin, et al.
Published: (2025) -
Low-Energy On-Device Personalization for MCUs
by: Huang, Yushan, et al.
Published: (2024) -
UnIT: Scalable Unstructured Inference-Time Pruning for MAC-efficient Neural Inference on MCUs
by: Neth, Ashe, et al.
Published: (2025) -
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
by: Zheng, Size, et al.
Published: (2024) -
ML Inference Scheduling with Predictable Latency
by: Zhao, Haidong, et al.
Published: (2025)