ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages
Fuente:
arXiv
Saved in:
| Main Author: | Qian, Zhoujie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CornViT: A Multi-Stage Convolutional Vision Transformer Framework for Hierarchical Corn Kernel Analysis
by: Erukude, Sai Teja, et al.
Published: (2025)
by: Erukude, Sai Teja, et al.
Published: (2025)
MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion
by: Hua, Wei, et al.
Published: (2025)
by: Hua, Wei, et al.
Published: (2025)
Spiking Vision Transformer with Saccadic Attention
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Revisiting the Integration of Convolution and Attention for Vision Backbone
by: Zhu, Lei, et al.
Published: (2024)
by: Zhu, Lei, et al.
Published: (2024)
HRPVT: High-Resolution Pyramid Vision Transformer for medium and small-scale human pose estimation
by: Xu, Zhoujie
Published: (2024)
by: Xu, Zhoujie
Published: (2024)
Advanced Brain Tumor Segmentation Using EMCAD: Efficient Multi-scale Convolutional Attention Decoding
by: Uzor, GodsGift, et al.
Published: (2025)
by: Uzor, GodsGift, et al.
Published: (2025)
An Efficient Dual-Line Decoder Network with Multi-Scale Convolutional Attention for Multi-organ Segmentation
by: Hassan, Riad, et al.
Published: (2025)
by: Hassan, Riad, et al.
Published: (2025)
Feature Fusion for Human Activity Recognition using Parameter-Optimized Multi-Stage Graph Convolutional Network and Transformer Models
by: Belal, Mohammad, et al.
Published: (2024)
by: Belal, Mohammad, et al.
Published: (2024)
The Linear Attention Resurrection in Vision Transformer
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
by: Zou, Dongyun, et al.
Published: (2026)
by: Zou, Dongyun, et al.
Published: (2026)
T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
by: Ntrougkas, Mariano V., et al.
Published: (2024)
by: Ntrougkas, Mariano V., et al.
Published: (2024)
Attention Retention for Continual Learning with Vision Transformers
by: Lu, Yue, et al.
Published: (2026)
by: Lu, Yue, et al.
Published: (2026)
Enhancing Satellite Object Localization with Dilated Convolutions and Attention-aided Spatial Pooling
by: Mostafa, Seraj Al Mahmud, et al.
Published: (2025)
by: Mostafa, Seraj Al Mahmud, et al.
Published: (2025)
Hierarchical Vision Transformer Enhanced by Graph Convolutional Network for Image Classification
by: Jiao, Haibin
Published: (2026)
by: Jiao, Haibin
Published: (2026)
Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-Attention
by: Leem, Saebom, et al.
Published: (2024)
by: Leem, Saebom, et al.
Published: (2024)
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
by: Böhle, Moritz, et al.
Published: (2025)
by: Böhle, Moritz, et al.
Published: (2025)
COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
by: Meng, Weikang, et al.
Published: (2025)
by: Meng, Weikang, et al.
Published: (2025)
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
by: Vani, Ankit, et al.
Published: (2024)
by: Vani, Ankit, et al.
Published: (2024)
Partial Convolution Meets Visual Attention
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
by: Tao, Hongyuan, et al.
Published: (2025)
by: Tao, Hongyuan, et al.
Published: (2025)
Feature Learning with Multi-Stage Vision Transformers on Inter-Modality HER2 Status Scoring and Tumor Classification on Whole Slides
by: Oyelade, Olaide N., et al.
Published: (2025)
by: Oyelade, Olaide N., et al.
Published: (2025)
LTMSformer: A Local Trend-Aware Attention and Motion State Encoding Transformer for Multi-Agent Trajectory Prediction
by: Yan, Yixin, et al.
Published: (2025)
by: Yan, Yixin, et al.
Published: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
by: Liang, Cheng, et al.
Published: (2026)
by: Liang, Cheng, et al.
Published: (2026)
Proximal Vision Transformer: Enhancing Feature Representation through Two-Stage Manifold Geometry
by: Yun, Haoyu, et al.
Published: (2025)
by: Yun, Haoyu, et al.
Published: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
Cross-Stage Attention Multi-Expert Network for Radiologist-Inspired Breast Ultrasound Diagnosis
by: Zhai, Xinyang, et al.
Published: (2026)
by: Zhai, Xinyang, et al.
Published: (2026)
Efficient Adaptation of Pre-trained Vision Transformer via Householder Transformation
by: Dong, Wei, et al.
Published: (2024)
by: Dong, Wei, et al.
Published: (2024)
Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
by: Carrión-Ojeda, Dustin, et al.
Published: (2025)
by: Carrión-Ojeda, Dustin, et al.
Published: (2025)
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
by: Apedo, Yvon, et al.
Published: (2026)
by: Apedo, Yvon, et al.
Published: (2026)
Heuristical Comparison of Vision Transformers Against Convolutional Neural Networks for Semantic Segmentation on Remote Sensing Imagery
by: Dahal, Ashim, et al.
Published: (2024)
by: Dahal, Ashim, et al.
Published: (2024)
Dynamic Accumulated Attention Map for Interpreting Evolution of Decision-Making in Vision Transformer
by: Liao, Yi, et al.
Published: (2025)
by: Liao, Yi, et al.
Published: (2025)
Symbolic Rule Extraction from Attention-Guided Sparse Representations in Vision Transformers
by: Padalkar, Parth, et al.
Published: (2025)
by: Padalkar, Parth, et al.
Published: (2025)
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
by: Knights, Ethan
Published: (2026)
by: Knights, Ethan
Published: (2026)
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025)
by: Lee, Dong-Jae, et al.
Published: (2025)
Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation
by: Cho, Yubin, et al.
Published: (2024)
by: Cho, Yubin, et al.
Published: (2024)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
CascadedViT: Cascaded Chunk-FeedForward and Cascaded Group Attention Vision Transformer
by: Sivakumar, Srivathsan, et al.
Published: (2025)
by: Sivakumar, Srivathsan, et al.
Published: (2025)
MUSTAN: Multi-scale Temporal Context as Attention for Robust Video Foreground Segmentation
by: Pokala, Praveen Kumar, et al.
Published: (2024)
by: Pokala, Praveen Kumar, et al.
Published: (2024)
Multi-scale Information Sharing and Selection Network with Boundary Attention for Polyp Segmentation
by: Kang, Xiaolu, et al.
Published: (2024)
by: Kang, Xiaolu, et al.
Published: (2024)
Similar Items
-
CornViT: A Multi-Stage Convolutional Vision Transformer Framework for Hierarchical Corn Kernel Analysis
by: Erukude, Sai Teja, et al.
Published: (2025) -
MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion
by: Hua, Wei, et al.
Published: (2025) -
Spiking Vision Transformer with Saccadic Attention
by: Wang, Shuai, et al.
Published: (2025) -
Revisiting the Integration of Convolution and Attention for Vision Backbone
by: Zhu, Lei, et al.
Published: (2024) -
HRPVT: High-Resolution Pyramid Vision Transformer for medium and small-scale human pose estimation
by: Xu, Zhoujie
Published: (2024)