Energy-Efficient Vision Transformer Inference for Edge-AI Deployment
Fuente:
arXiv
Saved in:
| Main Authors: | Amanzhol, Nursultan, Park, Jurn-Gyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient-Husformer: Efficient Multimodal Transformer Hyperparameter Optimization for Stress and Cognitive Loads
by: Orazaly, Merey, et al.
Published: (2025)
by: Orazaly, Merey, et al.
Published: (2025)
EdgeFlex-Transformer: Transformer Inference for Edge Devices
by: Mohammad, Shoaib, et al.
Published: (2025)
by: Mohammad, Shoaib, et al.
Published: (2025)
Resource-Efficient Generative AI Model Deployment in Mobile Edge Networks
by: Liang, Yuxin, et al.
Published: (2024)
by: Liang, Yuxin, et al.
Published: (2024)
LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
by: Nag, Shashank, et al.
Published: (2025)
by: Nag, Shashank, et al.
Published: (2025)
SparseDVFS: Sparse-Aware DVFS for Energy-Efficient Edge Inference
by: Zhang, Ziyang, et al.
Published: (2026)
by: Zhang, Ziyang, et al.
Published: (2026)
Co-Designing Binarized Transformer and Hardware Accelerator for Efficient End-to-End Edge Deployment
by: Ji, Yuhao, et al.
Published: (2024)
by: Ji, Yuhao, et al.
Published: (2024)
Adaptive Soft Rolling KV Freeze with Entropy-Guided Recovery: Sublinear Memory Growth for Efficient LLM Inference
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
The Larger the Merrier? Efficient Large AI Model Inference in Wireless Edge Networks
by: Lyu, Zhonghao, et al.
Published: (2025)
by: Lyu, Zhonghao, et al.
Published: (2025)
Transferable Deployment of Semantic Edge Inference Systems via Unsupervised Domain Adaption
by: Jiao, Weiqiang, et al.
Published: (2025)
by: Jiao, Weiqiang, et al.
Published: (2025)
Constrained Edge AI Deployment: Fine-Tuning vs Distillation for LLM Compression
by: Sander, Jacob, et al.
Published: (2025)
by: Sander, Jacob, et al.
Published: (2025)
AutoTailor: Automatic and Efficient Adaptive Model Deployment for Diverse Edge Devices
by: Liu, Mengyang, et al.
Published: (2025)
by: Liu, Mengyang, et al.
Published: (2025)
On the Sustainability of AI Inferences in the Edge
by: Sobhani, Ghazal, et al.
Published: (2025)
by: Sobhani, Ghazal, et al.
Published: (2025)
Graph Neural Networks Automated Design and Deployment on Device-Edge Co-Inference Systems
by: Zhou, Ao, et al.
Published: (2024)
by: Zhou, Ao, et al.
Published: (2024)
Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
by: Ye, Shengyuan, et al.
Published: (2024)
by: Ye, Shengyuan, et al.
Published: (2024)
QMC: Efficient SLM Edge Inference via Outlier-Aware Quantization and Emergent Memories Co-Design
by: Pandey, Nilesh Prasad, et al.
Published: (2026)
by: Pandey, Nilesh Prasad, et al.
Published: (2026)
Continuous-Time Homeostatic Dynamics for Reentrant Inference Models
by: Chae, Byung Gyu
Published: (2025)
by: Chae, Byung Gyu
Published: (2025)
Towards Efficient Deployment of Hybrid SNNs on Neuromorphic and Edge AI Hardware
by: Seekings, James, et al.
Published: (2024)
by: Seekings, James, et al.
Published: (2024)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
ElasticAI: Creating and Deploying Energy-Efficient Deep Learning Accelerator for Pervasive Computing
by: Qian, Chao, et al.
Published: (2024)
by: Qian, Chao, et al.
Published: (2024)
Ask the Expert: Collaborative Inference for Vision Transformers with Near-Edge Accelerators
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
On-Device Vision Training, Deployment, and Inference on a Thumb-Sized Microcontroller
by: Ellis, Jeremy
Published: (2026)
by: Ellis, Jeremy
Published: (2026)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
by: Kermani, Arshia, et al.
Published: (2025)
by: Kermani, Arshia, et al.
Published: (2025)
xTern: Energy-Efficient Ternary Neural Network Inference on RISC-V-Based Edge Systems
by: Rutishauser, Georg, et al.
Published: (2024)
by: Rutishauser, Georg, et al.
Published: (2024)
TinyFormer: Efficient Transformer Design and Deployment on Tiny Devices
by: Yang, Jianlei, et al.
Published: (2023)
by: Yang, Jianlei, et al.
Published: (2023)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
A Survey on Approximate Edge AI for Energy Efficient Autonomous Driving Services
by: Katare, Dewant, et al.
Published: (2023)
by: Katare, Dewant, et al.
Published: (2023)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
Adaptive Semantic Token Communication for Transformer-based Edge Inference
by: Devoto, Alessio, et al.
Published: (2025)
by: Devoto, Alessio, et al.
Published: (2025)
Efficient Autoregressive Inference for Transformer Probabilistic Models
by: Hassan, Conor, et al.
Published: (2025)
by: Hassan, Conor, et al.
Published: (2025)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
by: Park, Seong-Gyu, et al.
Published: (2026)
by: Park, Seong-Gyu, et al.
Published: (2026)
Sparse Optimization for Green Edge AI Inference
by: Yang, Xiangyu, et al.
Published: (2020)
by: Yang, Xiangyu, et al.
Published: (2020)
SPARQ: Spiking Early-Exit Neural Networks for Energy-Efficient Edge AI
by: Patne, Parth, et al.
Published: (2026)
by: Patne, Parth, et al.
Published: (2026)
ToFe: Lagged Token Freezing and Reusing for Efficient Vision Transformer Inference
by: Zhang, Haoyue, et al.
Published: (2025)
by: Zhang, Haoyue, et al.
Published: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
by: Husom, Erik Johannes, et al.
Published: (2025)
by: Husom, Erik Johannes, et al.
Published: (2025)
IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference
by: Zhong, Wanli, et al.
Published: (2025)
by: Zhong, Wanli, et al.
Published: (2025)
Rethinking the Use of Vision Transformers for AI-Generated Image Detection
by: Park, NaHyeon, et al.
Published: (2025)
by: Park, NaHyeon, et al.
Published: (2025)
Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
by: Ren, Ruifeng, et al.
Published: (2025)
by: Ren, Ruifeng, et al.
Published: (2025)
Shrinking the Giant : Quasi-Weightless Transformers for Low Energy Inference
by: Nag, Shashank, et al.
Published: (2024)
by: Nag, Shashank, et al.
Published: (2024)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
by: Yan, Minghao, et al.
Published: (2023)
by: Yan, Minghao, et al.
Published: (2023)
Similar Items
-
Efficient-Husformer: Efficient Multimodal Transformer Hyperparameter Optimization for Stress and Cognitive Loads
by: Orazaly, Merey, et al.
Published: (2025) -
EdgeFlex-Transformer: Transformer Inference for Edge Devices
by: Mohammad, Shoaib, et al.
Published: (2025) -
Resource-Efficient Generative AI Model Deployment in Mobile Edge Networks
by: Liang, Yuxin, et al.
Published: (2024) -
LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
by: Nag, Shashank, et al.
Published: (2025) -
SparseDVFS: Sparse-Aware DVFS for Energy-Efficient Edge Inference
by: Zhang, Ziyang, et al.
Published: (2026)