Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
Fuente:
arXiv
Salvato in:
| Autori principali: | Saha, Shaibal, Xu, Lanyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
di: Du, Dayou, et al.
Pubblicazione: (2024)
di: Du, Dayou, et al.
Pubblicazione: (2024)
EfficientQuant: An Efficient Post-Training Quantization for CNN-Transformer Hybrid Models on Edge Devices
di: Saha, Shaibal, et al.
Pubblicazione: (2025)
di: Saha, Shaibal, et al.
Pubblicazione: (2025)
SpNeRF: Memory Efficient Sparse Volumetric Neural Rendering Accelerator for Edge Devices
di: Zhang, Yipu, et al.
Pubblicazione: (2025)
di: Zhang, Yipu, et al.
Pubblicazione: (2025)
MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
di: Li, Shuaiting, et al.
Pubblicazione: (2024)
di: Li, Shuaiting, et al.
Pubblicazione: (2024)
Vision Transformer Computation and Resilience for Dynamic Inference
di: Sreedhar, Kavya, et al.
Pubblicazione: (2022)
di: Sreedhar, Kavya, et al.
Pubblicazione: (2022)
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design
di: You, Haoran, et al.
Pubblicazione: (2022)
di: You, Haoran, et al.
Pubblicazione: (2022)
Accelerating AI and Computer Vision for Satellite Pose Estimation on the Intel Myriad X Embedded SoC
di: Leon, Vasileios, et al.
Pubblicazione: (2024)
di: Leon, Vasileios, et al.
Pubblicazione: (2024)
LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge Devices
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
SF-MMCN: Low-Power Sever Flow Multi-Mode Diffusion Model Accelerator
di: Hsu, Huan-Ke, et al.
Pubblicazione: (2024)
di: Hsu, Huan-Ke, et al.
Pubblicazione: (2024)
A Parameterizable Convolution Accelerator for Embedded Deep Learning Applications
di: Mousouliotis, Panagiotis, et al.
Pubblicazione: (2026)
di: Mousouliotis, Panagiotis, et al.
Pubblicazione: (2026)
Hardware-Algorithm Co-Optimization of Early-Exit Neural Networks for Multi-Core Edge Accelerators
di: Zniber, Alaa, et al.
Pubblicazione: (2025)
di: Zniber, Alaa, et al.
Pubblicazione: (2025)
Real-Time Object Detection and Classification using YOLO for Edge FPGAs
di: Amin, Rashed Al, et al.
Pubblicazione: (2025)
di: Amin, Rashed Al, et al.
Pubblicazione: (2025)
hARMS: A Hardware Acceleration Architecture for Real-Time Event-Based Optical Flow
di: Stumpp, Daniel C., et al.
Pubblicazione: (2021)
di: Stumpp, Daniel C., et al.
Pubblicazione: (2021)
Primitive-Driven Acceleration of Hyperdimensional Computing for Real-Time Image Classification
di: Parikh, Dhruv, et al.
Pubblicazione: (2026)
di: Parikh, Dhruv, et al.
Pubblicazione: (2026)
TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
di: Miao, Wenxuan, et al.
Pubblicazione: (2025)
di: Miao, Wenxuan, et al.
Pubblicazione: (2025)
ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration
di: Lee, Hangyeol, et al.
Pubblicazione: (2026)
di: Lee, Hangyeol, et al.
Pubblicazione: (2026)
DPU or GPU for Accelerating Neural Networks Inference -- Why not both? Split CNN Inference
di: Oztas, Ali Emre, et al.
Pubblicazione: (2026)
di: Oztas, Ali Emre, et al.
Pubblicazione: (2026)
Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting Acceleration
di: Oh, Changhun, et al.
Pubblicazione: (2025)
di: Oh, Changhun, et al.
Pubblicazione: (2025)
EvGNN: An Event-driven Graph Neural Network Accelerator for Edge Vision
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
GLANCE: Gaze-Led Attention Network for Compressed Edge-inference
di: Solanki, Neeraj, et al.
Pubblicazione: (2026)
di: Solanki, Neeraj, et al.
Pubblicazione: (2026)
FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers
di: Durvasula, Sankeerth, et al.
Pubblicazione: (2025)
di: Durvasula, Sankeerth, et al.
Pubblicazione: (2025)
GS-TG: 3D Gaussian Splatting Accelerator with Tile Grouping for Reducing Redundant Sorting while Preserving Rasterization Efficiency
di: Jo, Joongho, et al.
Pubblicazione: (2025)
di: Jo, Joongho, et al.
Pubblicazione: (2025)
Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation
di: Carrigg, Kieran, et al.
Pubblicazione: (2026)
di: Carrigg, Kieran, et al.
Pubblicazione: (2026)
CHOSEN: Compilation to Hardware Optimization Stack for Efficient Vision Transformer Inference
di: Sadeghi, Mohammad Erfan, et al.
Pubblicazione: (2024)
di: Sadeghi, Mohammad Erfan, et al.
Pubblicazione: (2024)
Uni-Render: A Unified Accelerator for Real-Time Rendering Across Diverse Neural Renderers
di: Li, Chaojian, et al.
Pubblicazione: (2025)
di: Li, Chaojian, et al.
Pubblicazione: (2025)
Ditto: Accelerating Diffusion Model via Temporal Value Similarity
di: Kim, Sungbin, et al.
Pubblicazione: (2025)
di: Kim, Sungbin, et al.
Pubblicazione: (2025)
Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots
di: Williams, Justin, et al.
Pubblicazione: (2025)
di: Williams, Justin, et al.
Pubblicazione: (2025)
ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA
di: Lyu, Shengzhe, et al.
Pubblicazione: (2026)
di: Lyu, Shengzhe, et al.
Pubblicazione: (2026)
VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA
di: Wickramasinghe, Sachini, et al.
Pubblicazione: (2024)
di: Wickramasinghe, Sachini, et al.
Pubblicazione: (2024)
Performance Analysis of Edge and In-Sensor AI Processors: A Comparative Review
di: Capogrosso, Luigi, et al.
Pubblicazione: (2026)
di: Capogrosso, Luigi, et al.
Pubblicazione: (2026)
BILLNET: A Binarized Conv3D-LSTM Network with Logic-gated residual architecture for hardware-efficient video inference
di: Nguyen, Van Thien, et al.
Pubblicazione: (2025)
di: Nguyen, Van Thien, et al.
Pubblicazione: (2025)
From a Lossless (~1.5:1) Compression Algorithm for Llama2 7B Weights to Variable Precision, Variable Range, Compressed Numeric Data Types for CNNs and LLMs
di: Liguori, Vincenzo
Pubblicazione: (2024)
di: Liguori, Vincenzo
Pubblicazione: (2024)
Stella Nera: A Differentiable Maddness-Based Hardware Accelerator for Efficient Approximate Matrix Multiplication
di: Schönleber, Jannis, et al.
Pubblicazione: (2023)
di: Schönleber, Jannis, et al.
Pubblicazione: (2023)
SteROI-D: System Design and Mapping for Stereo Depth Inference on Regions of Interest
di: Erhardt, Jack, et al.
Pubblicazione: (2025)
di: Erhardt, Jack, et al.
Pubblicazione: (2025)
Using GUI Agent for Electronic Design Automation
di: Li, Chunyi, et al.
Pubblicazione: (2025)
di: Li, Chunyi, et al.
Pubblicazione: (2025)
Dedicated Inference Engine and Binary-Weight Neural Networks for Lightweight Instance Segmentation
di: Chen, Tse-Wei, et al.
Pubblicazione: (2025)
di: Chen, Tse-Wei, et al.
Pubblicazione: (2025)
AppSign: Multi-level Approximate Computing for Real-Time Traffic Sign Recognition in Autonomous Vehicles
di: Omidian, Fatemeh, et al.
Pubblicazione: (2024)
di: Omidian, Fatemeh, et al.
Pubblicazione: (2024)
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
di: Zhang, Baoheng, et al.
Pubblicazione: (2024)
di: Zhang, Baoheng, et al.
Pubblicazione: (2024)
Identifying Unnecessary 3D Gaussians using Clustering for Fast Rendering of 3D Gaussian Splatting
di: Jo, Joongho, et al.
Pubblicazione: (2024)
di: Jo, Joongho, et al.
Pubblicazione: (2024)
Gen-NeRF: Efficient and Generalizable Neural Radiance Fields via Algorithm-Hardware Co-Design
di: Fu, Yonggan, et al.
Pubblicazione: (2023)
di: Fu, Yonggan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
di: Du, Dayou, et al.
Pubblicazione: (2024) -
EfficientQuant: An Efficient Post-Training Quantization for CNN-Transformer Hybrid Models on Edge Devices
di: Saha, Shaibal, et al.
Pubblicazione: (2025) -
SpNeRF: Memory Efficient Sparse Volumetric Neural Rendering Accelerator for Edge Devices
di: Zhang, Yipu, et al.
Pubblicazione: (2025) -
MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
di: Li, Shuaiting, et al.
Pubblicazione: (2024) -
Vision Transformer Computation and Resilience for Dynamic Inference
di: Sreedhar, Kavya, et al.
Pubblicazione: (2022)