bViT: Investigating Single-Block Recurrence in Vision Transformers for Image Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Byra, Michal, Olszowiec, Pawel, Stefanski, Grzegorz, Gruszczynski, Grzegorz, Presta, Alberto |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Routing the Lottery: Adaptive Subnetworks for Heterogeneous Data
by: Stefanski, Grzegorz, et al.
Published: (2026)
by: Stefanski, Grzegorz, et al.
Published: (2026)
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition
by: Hu, Youbing, et al.
Published: (2024)
by: Hu, Youbing, et al.
Published: (2024)
t-gems: text-guided exit modules for decreasing clip image encoder
by: Presta, Alberto, et al.
Published: (2026)
by: Presta, Alberto, et al.
Published: (2026)
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025)
by: Jacobs, Mozes, et al.
Published: (2025)
Implantable Adaptive Cells: A Novel Enhancement for Pre-Trained U-Nets in Medical Image Segmentation
by: Benedykciuk, Emil, et al.
Published: (2024)
by: Benedykciuk, Emil, et al.
Published: (2024)
Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers with Bridge Block Reconstruction for IoT Systems
by: Lee, Jemin, et al.
Published: (2023)
by: Lee, Jemin, et al.
Published: (2023)
Efficient Search of Implantable Adaptive Cells for Medical Image Segmentation
by: Benedykciuk, Emil, et al.
Published: (2026)
by: Benedykciuk, Emil, et al.
Published: (2026)
LoViT: Long Video Transformer for Surgical Phase Recognition
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
VariViT: A Vision Transformer for Variable Image Sizes
by: Varma, Aswathi, et al.
Published: (2026)
by: Varma, Aswathi, et al.
Published: (2026)
Case-Enhanced Vision Transformer: Improving Explanations of Image Similarity with a ViT-based Similarity Metric
by: Zhao, Ziwei, et al.
Published: (2024)
by: Zhao, Ziwei, et al.
Published: (2024)
SAC-ViT: Semantic-Aware Clustering Vision Transformer with Early Exit
by: Hu, Youbing, et al.
Published: (2025)
by: Hu, Youbing, et al.
Published: (2025)
ButterflyViT: 354$\times$ Expert Compression for Edge Vision Transformers
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
MDS-ViTNet: Improving saliency prediction for Eye-Tracking with Vision Transformer
by: Ignat, Polezhaev, et al.
Published: (2024)
by: Ignat, Polezhaev, et al.
Published: (2024)
DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
by: Tong, Yujia, et al.
Published: (2025)
by: Tong, Yujia, et al.
Published: (2025)
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
by: Xu, Xuwei, et al.
Published: (2023)
by: Xu, Xuwei, et al.
Published: (2023)
ProtoPFormer: Concentrating on Prototypical Parts in Vision Transformers for Interpretable Image Recognition
by: Xue, Mengqi, et al.
Published: (2022)
by: Xue, Mengqi, et al.
Published: (2022)
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition
by: Alhadidi, Taqwa, et al.
Published: (2024)
by: Alhadidi, Taqwa, et al.
Published: (2024)
Unsupervised Tomato Split Anomaly Detection using Hyperspectral Imaging and Variational Autoencoders
by: Abdulsalam, Mahmoud, et al.
Published: (2025)
by: Abdulsalam, Mahmoud, et al.
Published: (2025)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
by: Zou, Dongyun, et al.
Published: (2026)
by: Zou, Dongyun, et al.
Published: (2026)
ViTs are Everywhere: A Comprehensive Study Showcasing Vision Transformers in Different Domain
by: Mia, Md Sohag, et al.
Published: (2023)
by: Mia, Md Sohag, et al.
Published: (2023)
CascadedViT: Cascaded Chunk-FeedForward and Cascaded Group Attention Vision Transformer
by: Sivakumar, Srivathsan, et al.
Published: (2025)
by: Sivakumar, Srivathsan, et al.
Published: (2025)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition
by: Kim, Ye-eun, et al.
Published: (2025)
by: Kim, Ye-eun, et al.
Published: (2025)
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
by: Góral, Gracjan, et al.
Published: (2025)
by: Góral, Gracjan, et al.
Published: (2025)
FasterViT: Fast Vision Transformers with Hierarchical Attention
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
by: Marmon, Andrew, et al.
Published: (2024)
by: Marmon, Andrew, et al.
Published: (2024)
Tiny-ViT: A Compact Vision Transformer for Efficient and Explainable Potato Leaf Disease Classification
by: Mia, Shakil, et al.
Published: (2026)
by: Mia, Shakil, et al.
Published: (2026)
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
by: Kim, Gihwan, et al.
Published: (2025)
by: Kim, Gihwan, et al.
Published: (2025)
CrisisViT: A Robust Vision Transformer for Crisis Image Classification
by: Long, Zijun, et al.
Published: (2024)
by: Long, Zijun, et al.
Published: (2024)
ScriptViT: Vision Transformer-Based Personalized Handwriting Generation
by: Acharya, Sajjan, et al.
Published: (2025)
by: Acharya, Sajjan, et al.
Published: (2025)
Octic Vision Transformers: Quicker ViTs Through Equivariance
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
CornViT: A Multi-Stage Convolutional Vision Transformer Framework for Hierarchical Corn Kernel Analysis
by: Erukude, Sai Teja, et al.
Published: (2025)
by: Erukude, Sai Teja, et al.
Published: (2025)
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
by: Xu, Xuwei, et al.
Published: (2025)
by: Xu, Xuwei, et al.
Published: (2025)
Generating visual explanations from deep networks using implicit neural representations
by: Byra, Michal, et al.
Published: (2025)
by: Byra, Michal, et al.
Published: (2025)
SugarViT -- Multi-objective Regression of UAV Images with Vision Transformers and Deep Label Distribution Learning Demonstrated on Disease Severity Prediction in Sugar Beet
by: Günder, Maurice, et al.
Published: (2023)
by: Günder, Maurice, et al.
Published: (2023)
Semantically-Prompted Language Models Improve Visual Descriptions
by: Ogezi, Michael, et al.
Published: (2023)
by: Ogezi, Michael, et al.
Published: (2023)
Generalized Single-Image-Based Morphing Attack Detection Using Deep Representations from Vision Transformer
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
Flexible ViG: Learning the Self-Saliency for Flexible Object Recognition
by: Zuo, Lin, et al.
Published: (2024)
by: Zuo, Lin, et al.
Published: (2024)
TinyViT-Batten: Few-Shot Vision Transformer with Explainable Attention for Early Batten-Disease Detection on Pediatric MRI
by: Uppalapati, Khartik, et al.
Published: (2025)
by: Uppalapati, Khartik, et al.
Published: (2025)
Sensitive Image Classification by Vision Transformers
by: He, Hanxian, et al.
Published: (2024)
by: He, Hanxian, et al.
Published: (2024)
Similar Items
-
Routing the Lottery: Adaptive Subnetworks for Heterogeneous Data
by: Stefanski, Grzegorz, et al.
Published: (2026) -
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition
by: Hu, Youbing, et al.
Published: (2024) -
t-gems: text-guided exit modules for decreasing clip image encoder
by: Presta, Alberto, et al.
Published: (2026) -
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025) -
Implantable Adaptive Cells: A Novel Enhancement for Pre-Trained U-Nets in Medical Image Segmentation
by: Benedykciuk, Emil, et al.
Published: (2024)