QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Oh, Sehyeon, Kwon, Yongin, Lee, Jemin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixed Non-linear Quantization for Vision Transformers
by: Kim, Gihwan, et al.
Published: (2024)
by: Kim, Gihwan, et al.
Published: (2024)
Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers with Bridge Block Reconstruction for IoT Systems
by: Lee, Jemin, et al.
Published: (2023)
by: Lee, Jemin, et al.
Published: (2023)
Exploring the Trade-Offs: Quantization Methods, Task Difficulty, and Model Size in Large Language Models From Edge to Giant
by: Lee, Jemin, et al.
Published: (2024)
by: Lee, Jemin, et al.
Published: (2024)
QuantuneV2: Compiler-Based Local Metric-Driven Mixed Precision Quantization for Practical Embedded AI Applications
by: Kim, Jeongseok, et al.
Published: (2025)
by: Kim, Jeongseok, et al.
Published: (2025)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
by: Kim, Taeho, et al.
Published: (2024)
by: Kim, Taeho, et al.
Published: (2024)
A Predictive Model Based on Transformer with Statistical Feature Embedding in Manufacturing Sensor Dataset
by: Lee, Gyeong Taek, et al.
Published: (2024)
by: Lee, Gyeong Taek, et al.
Published: (2024)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
by: Choi, Kanghyun, et al.
Published: (2024)
by: Choi, Kanghyun, et al.
Published: (2024)
ML$^2$Tuner: Efficient Code Tuning via Multi-Level Machine Learning Models
by: Cha, JooHyoung, et al.
Published: (2024)
by: Cha, JooHyoung, et al.
Published: (2024)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
Echo State Transformer: Attention Over Finite Memories
by: Bendi-Ouis, Yannis, et al.
Published: (2025)
by: Bendi-Ouis, Yannis, et al.
Published: (2025)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)
by: Lee, Jung Hyun, et al.
Published: (2023)
LOOKAT: Lookup-Optimized Key-Attention for Memory-Efficient Transformers
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
Dispatch-Aware Ragged Attention for Pruned Vision Transformers
by: Abdellatif, Seifeldin, et al.
Published: (2026)
by: Abdellatif, Seifeldin, et al.
Published: (2026)
Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning
by: Lee, Joonho, et al.
Published: (2019)
by: Lee, Joonho, et al.
Published: (2019)
BoA: Attention-aware Post-training Quantization without Backpropagation
by: Kim, Junhan, et al.
Published: (2024)
by: Kim, Junhan, et al.
Published: (2024)
Not Only Rewards But Also Constraints: Applications on Legged Robot Locomotion
by: Kim, Yunho, et al.
Published: (2023)
by: Kim, Yunho, et al.
Published: (2023)
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
by: Deiseroth, Björn, et al.
Published: (2023)
by: Deiseroth, Björn, et al.
Published: (2023)
Enhancing Training Efficiency Using Packing with Flash Attention
by: Kundu, Achintya, et al.
Published: (2024)
by: Kundu, Achintya, et al.
Published: (2024)
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
by: Kim, Gihwan, et al.
Published: (2025)
by: Kim, Gihwan, et al.
Published: (2025)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
by: Yang, June Yong, et al.
Published: (2024)
by: Yang, June Yong, et al.
Published: (2024)
Grouped Differential Attention
by: Lim, Junghwan, et al.
Published: (2025)
by: Lim, Junghwan, et al.
Published: (2025)
QSViT: A Methodology for Quantizing Spiking Vision Transformers
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2025)
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2025)
LoopQ: Quantization for Recursive Transformers
by: Fang, Rui, et al.
Published: (2026)
by: Fang, Rui, et al.
Published: (2026)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
by: Kim, Junhan, et al.
Published: (2024)
by: Kim, Junhan, et al.
Published: (2024)
Graph Convolutions Enrich the Self-Attention in Transformers!
by: Choi, Jeongwhan, et al.
Published: (2023)
by: Choi, Jeongwhan, et al.
Published: (2023)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
by: Zhang, Peiyuan, et al.
Published: (2026)
by: Zhang, Peiyuan, et al.
Published: (2026)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
by: Lee, Jung Hyun, et al.
Published: (2024)
by: Lee, Jung Hyun, et al.
Published: (2024)
Parameter Efficiency Is Not Memory Efficiency: Rethinking Fine-Tuning for On-Device LLM Adaptation
by: Tenison, Irene, et al.
Published: (2026)
by: Tenison, Irene, et al.
Published: (2026)
Sobolev acceleration for neural networks
by: Oh, Jong Kwon, et al.
Published: (2025)
by: Oh, Jong Kwon, et al.
Published: (2025)
Scratching Visual Transformer's Back with Uniform Attention
by: Hyeon-Woo, Nam, et al.
Published: (2022)
by: Hyeon-Woo, Nam, et al.
Published: (2022)
PolarQuant: Quantizing KV Caches with Polar Transformation
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
Adaptive Memory Decay for Log-Linear Attention
by: Amin, Yaxita, et al.
Published: (2026)
by: Amin, Yaxita, et al.
Published: (2026)
Fusing Memory and Attention: A study on LSTM, Transformer and Hybrid Architectures for Symbolic Music Generation
by: Ghoshal, Soudeep, et al.
Published: (2026)
by: Ghoshal, Soudeep, et al.
Published: (2026)
Graph Tokenization for Bridging Graphs and Transformers
by: Guo, Zeyuan, et al.
Published: (2026)
by: Guo, Zeyuan, et al.
Published: (2026)
Paged Attention Meets FlexAttention: Unlocking Long-Context Efficiency in Deployed Inference
by: Joshi, Thomas, et al.
Published: (2025)
by: Joshi, Thomas, et al.
Published: (2025)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
by: You, Jaeseong, et al.
Published: (2024)
by: You, Jaeseong, et al.
Published: (2024)
Recurrent Action Transformer with Memory
by: Cherepanov, Egor, et al.
Published: (2023)
by: Cherepanov, Egor, et al.
Published: (2023)
Similar Items
-
Mixed Non-linear Quantization for Vision Transformers
by: Kim, Gihwan, et al.
Published: (2024) -
Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers with Bridge Block Reconstruction for IoT Systems
by: Lee, Jemin, et al.
Published: (2023) -
Exploring the Trade-Offs: Quantization Methods, Task Difficulty, and Model Size in Large Language Models From Edge to Giant
by: Lee, Jemin, et al.
Published: (2024) -
QuantuneV2: Compiler-Based Local Metric-Driven Mixed Precision Quantization for Practical Embedded AI Applications
by: Kim, Jeongseok, et al.
Published: (2025) -
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
by: Kim, Taeho, et al.
Published: (2024)