SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Jialong, Chen, Xinghao, Tang, Yehui, Wang, Yunhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation
by: Ni, Zhenliang, et al.
Published: (2024)
by: Ni, Zhenliang, et al.
Published: (2024)
A Survey on Transformer Compression
by: Tang, Yehui, et al.
Published: (2024)
by: Tang, Yehui, et al.
Published: (2024)
No Time to Waste: Squeeze Time into Channel for Mobile Video Understanding
by: Zhai, Yingjie, et al.
Published: (2024)
by: Zhai, Yingjie, et al.
Published: (2024)
Token Compensator: Altering Inference Cost of Vision Transformer without Re-Tuning
by: Jie, Shibo, et al.
Published: (2024)
by: Jie, Shibo, et al.
Published: (2024)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
by: Jie, Shibo, et al.
Published: (2024)
by: Jie, Shibo, et al.
Published: (2024)
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
by: Shu, Han, et al.
Published: (2023)
by: Shu, Han, et al.
Published: (2023)
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
by: Ni, Zhenliang, et al.
Published: (2025)
by: Ni, Zhenliang, et al.
Published: (2025)
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
by: Rang, Miao, et al.
Published: (2025)
by: Rang, Miao, et al.
Published: (2025)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
by: Han, Kai, et al.
Published: (2024)
by: Han, Kai, et al.
Published: (2024)
RepGhost: A Hardware-Efficient Ghost Module via Re-parameterization
by: Chen, Chengpeng, et al.
Published: (2022)
by: Chen, Chengpeng, et al.
Published: (2022)
GhostNetV3: Exploring the Training Strategies for Compact Models
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
DECO: Unleashing the Potential of ConvNets for Query-based Detection and Segmentation
by: Chen, Xinghao, et al.
Published: (2023)
by: Chen, Xinghao, et al.
Published: (2023)
LoReTrack: Efficient and Accurate Low-Resolution Transformer Tracking
by: Dong, Shaohua, et al.
Published: (2024)
by: Dong, Shaohua, et al.
Published: (2024)
IPT-V2: Efficient Image Processing Transformer using Hierarchical Attentions
by: Tu, Zhijun, et al.
Published: (2024)
by: Tu, Zhijun, et al.
Published: (2024)
Prism: Spectral-Aware Block-Sparse Attention
by: Wang, Xinghao, et al.
Published: (2026)
by: Wang, Xinghao, et al.
Published: (2026)
SAM-DiffSR: Structure-Modulated Diffusion Model for Image Super-Resolution
by: Wang, Chengcheng, et al.
Published: (2024)
by: Wang, Chengcheng, et al.
Published: (2024)
ASR: Attention-alike Structural Re-parameterization
by: Zhong, Shanshan, et al.
Published: (2023)
by: Zhong, Shanshan, et al.
Published: (2023)
Multi-Scale Correlation-Aware Transformer for Maritime Vessel Re-Identification
by: Liu, Yunhe
Published: (2025)
by: Liu, Yunhe
Published: (2025)
Batch Transformer: Look for Attention in Batch
by: Her, Myung Beom, et al.
Published: (2024)
by: Her, Myung Beom, et al.
Published: (2024)
GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task
by: Ding, Ning, et al.
Published: (2023)
by: Ding, Ning, et al.
Published: (2023)
PPT: Token Pruning and Pooling for Efficient Vision Transformers
by: Wu, Xinjian, et al.
Published: (2023)
by: Wu, Xinjian, et al.
Published: (2023)
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
by: Kang, Jialiang, et al.
Published: (2025)
by: Kang, Jialiang, et al.
Published: (2025)
Till the Layers Collapse: Compressing a Deep Neural Network through the Lenses of Batch Normalization Layers
by: Liao, Zhu, et al.
Published: (2024)
by: Liao, Zhu, et al.
Published: (2024)
Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning
by: Chen, Yanjun, et al.
Published: (2025)
by: Chen, Yanjun, et al.
Published: (2025)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
by: Mao, Weian, et al.
Published: (2026)
by: Mao, Weian, et al.
Published: (2026)
Sparser Block-Sparse Attention via Token Permutation
by: Wang, Xinghao, et al.
Published: (2025)
by: Wang, Xinghao, et al.
Published: (2025)
Multi-Scale Implicit Transformer with Re-parameterize for Arbitrary-Scale Super-Resolution
by: Zhu, Jinchen, et al.
Published: (2024)
by: Zhu, Jinchen, et al.
Published: (2024)
Transformers without Normalization
by: Zhu, Jiachen, et al.
Published: (2025)
by: Zhu, Jiachen, et al.
Published: (2025)
UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
by: Tang, Hongxuan, et al.
Published: (2025)
by: Tang, Hongxuan, et al.
Published: (2025)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
by: Liu, Wenjie, et al.
Published: (2026)
by: Liu, Wenjie, et al.
Published: (2026)
Normalizing Batch Normalization for Long-Tailed Recognition
by: Bao, Yuxiang, et al.
Published: (2025)
by: Bao, Yuxiang, et al.
Published: (2025)
UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
Stronger Normalization-Free Transformers
by: Chen, Mingzhi, et al.
Published: (2025)
by: Chen, Mingzhi, et al.
Published: (2025)
Boosting Pruned Networks with Linear Over-parameterization
by: Qian, Yu, et al.
Published: (2022)
by: Qian, Yu, et al.
Published: (2022)
Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
by: Naz, Zubia, et al.
Published: (2025)
by: Naz, Zubia, et al.
Published: (2025)
Progressively Normalized Self-Attention Network for Video Polyp Segmentation
by: Ji, Ge-Peng, et al.
Published: (2021)
by: Ji, Ge-Peng, et al.
Published: (2021)
DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
by: Qi, Xianbiao, et al.
Published: (2025)
by: Qi, Xianbiao, et al.
Published: (2025)
Spatial Re-parameterization for N:M Sparsity
by: Zhang, Yuxin, et al.
Published: (2023)
by: Zhang, Yuxin, et al.
Published: (2023)
VAEER: Visual Attention-Inspired Emotion Elicitation Reasoning
by: Man, Fanhang, et al.
Published: (2025)
by: Man, Fanhang, et al.
Published: (2025)
ReGLA: Efficient Receptive-Field Modeling with Gated Linear Attention Network
by: Li, Junzhou, et al.
Published: (2026)
by: Li, Junzhou, et al.
Published: (2026)
Similar Items
-
Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation
by: Ni, Zhenliang, et al.
Published: (2024) -
A Survey on Transformer Compression
by: Tang, Yehui, et al.
Published: (2024) -
No Time to Waste: Squeeze Time into Channel for Mobile Video Understanding
by: Zhai, Yingjie, et al.
Published: (2024) -
Token Compensator: Altering Inference Cost of Vision Transformer without Re-Tuning
by: Jie, Shibo, et al.
Published: (2024) -
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
by: Jie, Shibo, et al.
Published: (2024)