Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Dayou, Gong, Gu, Chu, Xiaowen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
by: Aggarwal, Shivam, et al.
Published: (2023)
by: Aggarwal, Shivam, et al.
Published: (2023)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity
by: Fan, Zichen, et al.
Published: (2025)
by: Fan, Zichen, et al.
Published: (2025)
Neural Architecture Search of Hybrid Models for NPU-CIM Heterogeneous AR/VR Devices
by: Zhao, Yiwei, et al.
Published: (2024)
by: Zhao, Yiwei, et al.
Published: (2024)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
by: Patwari, Rajeev, et al.
Published: (2025)
by: Patwari, Rajeev, et al.
Published: (2025)
On Latency Predictors for Neural Architecture Search
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
by: Chhugani, Jatin, et al.
Published: (2026)
by: Chhugani, Jatin, et al.
Published: (2026)
CHOSEN: Compilation to Hardware Optimization Stack for Efficient Vision Transformer Inference
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
by: Lübeck, Konstantin, et al.
Published: (2024)
by: Lübeck, Konstantin, et al.
Published: (2024)
Design Insights and Comparative Evaluation of a Hardware-Based Cooperative Perception Architecture for Lane Change Prediction
by: Manzour, Mohamed, et al.
Published: (2025)
by: Manzour, Mohamed, et al.
Published: (2025)
Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
by: Saha, Shaibal, et al.
Published: (2025)
by: Saha, Shaibal, et al.
Published: (2025)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
by: Müller, Mika Markus, et al.
Published: (2025)
by: Müller, Mika Markus, et al.
Published: (2025)
AHCQ-SAM: Toward Accurate and Hardware-Compatible Post-Training Segment Anything Model Quantization
by: Zhang, Wenlun, et al.
Published: (2025)
by: Zhang, Wenlun, et al.
Published: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025)
by: Du, Dayou, et al.
Published: (2025)
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA
by: Lyu, Shengzhe, et al.
Published: (2026)
by: Lyu, Shengzhe, et al.
Published: (2026)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
by: Zhang, Hang, et al.
Published: (2025)
by: Zhang, Hang, et al.
Published: (2025)
Stella Nera: A Differentiable Maddness-Based Hardware Accelerator for Efficient Approximate Matrix Multiplication
by: Schönleber, Jannis, et al.
Published: (2023)
by: Schönleber, Jannis, et al.
Published: (2023)
FabGPT: An Efficient Large Multimodal Model for Complex Wafer Defect Knowledge Queries
by: Jiang, Yuqi, et al.
Published: (2024)
by: Jiang, Yuqi, et al.
Published: (2024)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
by: Saha, Rappy, et al.
Published: (2026)
by: Saha, Rappy, et al.
Published: (2026)
CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
KWT-Tiny: RISC-V Accelerated, Embedded Keyword Spotting Transformer
by: Al-Qawlaq, Aness, et al.
Published: (2024)
by: Al-Qawlaq, Aness, et al.
Published: (2024)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
by: Yang, Hanchen, et al.
Published: (2025)
by: Yang, Hanchen, et al.
Published: (2025)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
by: Bi, Zhen, et al.
Published: (2026)
by: Bi, Zhen, et al.
Published: (2026)
Characterizing and Understanding HGNN Training on GPUs
by: Han, Dengke, et al.
Published: (2024)
by: Han, Dengke, et al.
Published: (2024)
Towards Efficient Deployment of Hybrid SNNs on Neuromorphic and Edge AI Hardware
by: Seekings, James, et al.
Published: (2024)
by: Seekings, James, et al.
Published: (2024)
Strassen Multisystolic Array Hardware Architectures
by: Pogue, Trevor E., et al.
Published: (2025)
by: Pogue, Trevor E., et al.
Published: (2025)
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
by: Shi, Tianyao, et al.
Published: (2025)
by: Shi, Tianyao, et al.
Published: (2025)
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
by: Odema, Mohanad, et al.
Published: (2024)
by: Odema, Mohanad, et al.
Published: (2024)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
by: Panigrahy, Deepak, et al.
Published: (2026)
by: Panigrahy, Deepak, et al.
Published: (2026)
Real-World Deployment of a Lane Change Prediction Architecture Based on Knowledge Graph Embeddings and Bayesian Inference
by: Manzour, M., et al.
Published: (2025)
by: Manzour, M., et al.
Published: (2025)
TSLA: A Task-Specific Learning Adaptation for Semantic Segmentation on Autonomous Vehicles Platform
by: Liu, Jun, et al.
Published: (2025)
by: Liu, Jun, et al.
Published: (2025)
fpgaHART: A toolflow for throughput-oriented acceleration of 3D CNNs for HAR onto FPGAs
by: Toupas, Petros, et al.
Published: (2023)
by: Toupas, Petros, et al.
Published: (2023)
SageAttention2++: A More Efficient Implementation of SageAttention2
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
Real-Time Semantic Segmentation of Aerial Images Using an Embedded U-Net: A Comparison of CPU, GPU, and FPGA Workflows
by: Posso, Julien, et al.
Published: (2025)
by: Posso, Julien, et al.
Published: (2025)
Rapid-INR: Storage Efficient CPU-free DNN Training Using Implicit Neural Representation
by: Chen, Hanqiu, et al.
Published: (2023)
by: Chen, Hanqiu, et al.
Published: (2023)
FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition
by: Toupas, Petros, et al.
Published: (2023)
by: Toupas, Petros, et al.
Published: (2023)
Similar Items
-
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
by: Aggarwal, Shivam, et al.
Published: (2023) -
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
by: Zhang, Jintao, et al.
Published: (2025) -
SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity
by: Fan, Zichen, et al.
Published: (2025) -
Neural Architecture Search of Hybrid Models for NPU-CIM Heterogeneous AR/VR Devices
by: Zhao, Yiwei, et al.
Published: (2024) -
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
by: Patwari, Rajeev, et al.
Published: (2025)