Revisiting the Integration of Convolution and Attention for Vision Backbone
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Lei, Wang, Xinjiang, Zhang, Wayne, Lau, Rynson W. H. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MirrorMamba: Towards Scalable and Robust Mirror Detection in Videos
by: Song, Rui, et al.
Published: (2025)
by: Song, Rui, et al.
Published: (2025)
RapidNet: Multi-Level Dilated Convolution Based Mobile Backbone
by: Munir, Mustafa, et al.
Published: (2024)
by: Munir, Mustafa, et al.
Published: (2024)
Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
by: Du, Yifan, et al.
Published: (2025)
by: Du, Yifan, et al.
Published: (2025)
A Generic Shared Attention Mechanism for Various Backbone Neural Networks
by: Huang, Zhongzhan, et al.
Published: (2022)
by: Huang, Zhongzhan, et al.
Published: (2022)
Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones
by: Nottebaum, Moritz, et al.
Published: (2026)
by: Nottebaum, Moritz, et al.
Published: (2026)
ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages
by: Qian, Zhoujie
Published: (2025)
by: Qian, Zhoujie
Published: (2025)
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
by: Liu, Haixu, et al.
Published: (2025)
by: Liu, Haixu, et al.
Published: (2025)
CPUBone: Efficient Vision Backbone Design for Devices with Low Parallelization Capabilities
by: Nottebaum, Moritz, et al.
Published: (2026)
by: Nottebaum, Moritz, et al.
Published: (2026)
Inverse Rendering of Glossy Objects via the Neural Plenoptic Function and Radiance Fields
by: Wang, Haoyuan, et al.
Published: (2024)
by: Wang, Haoyuan, et al.
Published: (2024)
Vision-LSTM: xLSTM as Generic Vision Backbone
by: Alkin, Benedikt, et al.
Published: (2024)
by: Alkin, Benedikt, et al.
Published: (2024)
MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
Partial Convolution Meets Visual Attention
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Spiking Vision Transformer with Saccadic Attention
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
ViR: Towards Efficient Vision Retention Backbones
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
Improved Belief-Attention in Vision Task
by: Zhang, Guoqiang
Published: (2026)
by: Zhang, Guoqiang
Published: (2026)
Enhancing Satellite Object Localization with Dilated Convolutions and Attention-aided Spatial Pooling
by: Mostafa, Seraj Al Mahmud, et al.
Published: (2025)
by: Mostafa, Seraj Al Mahmud, et al.
Published: (2025)
Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation
by: Liu, Xuewen, et al.
Published: (2025)
by: Liu, Xuewen, et al.
Published: (2025)
Attention Retention for Continual Learning with Vision Transformers
by: Lu, Yue, et al.
Published: (2026)
by: Lu, Yue, et al.
Published: (2026)
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
by: Barsellotti, Luca, et al.
Published: (2024)
by: Barsellotti, Luca, et al.
Published: (2024)
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
by: Horawalavithana, Sameera, et al.
Published: (2026)
by: Horawalavithana, Sameera, et al.
Published: (2026)
Revisiting Cross-Attention Mechanisms: Leveraging Beneficial Noise for Domain-Adaptive Learning
by: Zang, Zelin, et al.
Published: (2026)
by: Zang, Zelin, et al.
Published: (2026)
Vision Backbone Efficient Selection for Image Classification in Low-Data Regimes
by: Guerin, Joris, et al.
Published: (2024)
by: Guerin, Joris, et al.
Published: (2024)
SAFE-KD: Risk-Controlled Early-Exit Distillation for Vision Backbones
by: Khazem, Salim
Published: (2026)
by: Khazem, Salim
Published: (2026)
A Survey on Backbones for Deep Video Action Recognition
by: Tang, Zixuan, et al.
Published: (2024)
by: Tang, Zixuan, et al.
Published: (2024)
Normalization Equivariance for Arbitrary Backbones, with Application to Image Denoising
by: Saied, Youssef, et al.
Published: (2026)
by: Saied, Youssef, et al.
Published: (2026)
Attention Prompting on Image for Large Vision-Language Models
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
Bladder Vessel Segmentation using a Hybrid Attention-Convolution Framework
by: Krauß, Franziska, et al.
Published: (2026)
by: Krauß, Franziska, et al.
Published: (2026)
Large Vision-Language Models Get Lost in Attention
by: Xi, Gongli, et al.
Published: (2026)
by: Xi, Gongli, et al.
Published: (2026)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
by: Zhu, Younan, et al.
Published: (2025)
by: Zhu, Younan, et al.
Published: (2025)
Generalized Large-Scale Data Condensation via Various Backbone and Statistical Matching
by: Shao, Shitong, et al.
Published: (2023)
by: Shao, Shitong, et al.
Published: (2023)
MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
by: Wang, Yuhan, et al.
Published: (2025)
by: Wang, Yuhan, et al.
Published: (2025)
T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
by: Ntrougkas, Mariano V., et al.
Published: (2024)
by: Ntrougkas, Mariano V., et al.
Published: (2024)
IGASA: Integrated Geometry-Aware and Skip-Attention Modules for Enhanced Point Cloud Registration
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
A-VL: Adaptive Attention for Large Vision-Language Models
by: Zhang, Junyang, et al.
Published: (2024)
by: Zhang, Junyang, et al.
Published: (2024)
The Linear Attention Resurrection in Vision Transformer
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
Self-Supervised Backbone Framework for Diverse Agricultural Vision Tasks
by: Sornapudi, Sudhir, et al.
Published: (2024)
by: Sornapudi, Sudhir, et al.
Published: (2024)
Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
by: Zhang, Naifu, et al.
Published: (2025)
by: Zhang, Naifu, et al.
Published: (2025)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
by: Zhang, Yudong, et al.
Published: (2024)
by: Zhang, Yudong, et al.
Published: (2024)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
Similar Items
-
MirrorMamba: Towards Scalable and Robust Mirror Detection in Videos
by: Song, Rui, et al.
Published: (2025) -
RapidNet: Multi-Level Dilated Convolution Based Mobile Backbone
by: Munir, Mustafa, et al.
Published: (2024) -
Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
by: Du, Yifan, et al.
Published: (2025) -
A Generic Shared Attention Mechanism for Various Backbone Neural Networks
by: Huang, Zhongzhan, et al.
Published: (2022) -
Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones
by: Nottebaum, Moritz, et al.
Published: (2026)