Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
Fuente:
arXiv
Saved in:
| Main Authors: | Yunusa, Haruna, Qin, Shiyin, Chukkol, Abdulrahman Hamman Adama, Yusuf, Abdulganiyu Abdu, Bello, Isah, Lawan, Adamu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SaRPFF: A Self-Attention with Register-based Pyramid Feature Fusion module for enhanced RLD detection
by: Haruna, Yunusa, et al.
Published: (2024)
by: Haruna, Yunusa, et al.
Published: (2024)
KonvLiNA: Integrating Kolmogorov-Arnold Network with Linear Nyström Attention for feature fusion in Crop Field Detection
by: Yunusa, Haruna, et al.
Published: (2024)
by: Yunusa, Haruna, et al.
Published: (2024)
iiANET: Inception Inspired Attention Hybrid Network for efficient Long-Range Dependency
by: Yunusa, Haruna, et al.
Published: (2024)
by: Yunusa, Haruna, et al.
Published: (2024)
VulCatch: Enhancing Binary Vulnerability Detection through CodeT5 Decompilation and KAN Advanced Feature Extraction
by: Chukkol, Abdulrahman Hamman Adama, et al.
Published: (2024)
by: Chukkol, Abdulrahman Hamman Adama, et al.
Published: (2024)
GateMABSA: Aspect-Image Gated Fusion for Multimodal Aspect-based Sentiment Analysis
by: Lawan, Adamu, et al.
Published: (2025)
by: Lawan, Adamu, et al.
Published: (2025)
Enhancing Long-Range Dependency with State Space Model and Kolmogorov-Arnold Networks for Aspect-Based Sentiment Analysis
by: Lawan, Adamu, et al.
Published: (2024)
by: Lawan, Adamu, et al.
Published: (2024)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
by: Shah, Arya, et al.
Published: (2025)
by: Shah, Arya, et al.
Published: (2025)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
by: Chen, Lu, et al.
Published: (2025)
by: Chen, Lu, et al.
Published: (2025)
AF-MAT: Aspect-aware Flip-and-Fuse xLSTM for Aspect-based Sentiment Analysis
by: Lawan, Adamu, et al.
Published: (2025)
by: Lawan, Adamu, et al.
Published: (2025)
DualKanbaFormer: An Efficient Selective Sparse Framework for Multimodal Aspect-based Sentiment Analysis
by: Lawan, Adamu, et al.
Published: (2024)
by: Lawan, Adamu, et al.
Published: (2024)
Amplifying Aspect-Sentence Awareness: A Novel Approach for Aspect-Based Sentiment Analysis
by: Lawan, Adamu, et al.
Published: (2024)
by: Lawan, Adamu, et al.
Published: (2024)
ViTs are Everywhere: A Comprehensive Study Showcasing Vision Transformers in Different Domain
by: Mia, Md Sohag, et al.
Published: (2023)
by: Mia, Md Sohag, et al.
Published: (2023)
vGamba: Attentive State Space Bottleneck for efficient Long-range Dependencies in Visual Recognition
by: Haruna, Yunusa, et al.
Published: (2025)
by: Haruna, Yunusa, et al.
Published: (2025)
Refining Datapath for Microscaling ViTs
by: Xiao, Can, et al.
Published: (2025)
by: Xiao, Can, et al.
Published: (2025)
Octic Vision Transformers: Quicker ViTs Through Equivariance
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
Bias Redistribution in Visual Machine Unlearning: Does Forgetting One Group Harm Another?
by: Haruna, Yunusa, et al.
Published: (2026)
by: Haruna, Yunusa, et al.
Published: (2026)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
by: Zhong, Yunshan, et al.
Published: (2023)
by: Zhong, Yunshan, et al.
Published: (2023)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
by: Ramachandran, Akshat, et al.
Published: (2024)
by: Ramachandran, Akshat, et al.
Published: (2024)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
by: Chattopadhyay, Nandish, et al.
Published: (2026)
by: Chattopadhyay, Nandish, et al.
Published: (2026)
TAP-ViTs: Task-Adaptive Pruning for On-Device Deployment of Vision Transformers
by: Wang, Zhibo, et al.
Published: (2026)
by: Wang, Zhibo, et al.
Published: (2026)
TransResNet: Integrating the Strengths of ViTs and CNNs for High Resolution Medical Image Segmentation via Feature Grafting
by: Sharif, Muhammad Hamza, et al.
Published: (2024)
by: Sharif, Muhammad Hamza, et al.
Published: (2024)
U-REPA: Aligning Diffusion U-Nets to ViTs
by: Tian, Yuchuan, et al.
Published: (2025)
by: Tian, Yuchuan, et al.
Published: (2025)
Elastic ViTs from Pretrained Models without Retraining
by: Simoncini, Walter, et al.
Published: (2025)
by: Simoncini, Walter, et al.
Published: (2025)
Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability
by: Liu, Jiani, et al.
Published: (2025)
by: Liu, Jiani, et al.
Published: (2025)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
Fusing Pretrained ViTs with TCNet for Enhanced EEG Regression
by: Modesitt, Eric, et al.
Published: (2024)
by: Modesitt, Eric, et al.
Published: (2024)
Pretrained ViTs Yield Versatile Representations For Medical Images
by: Matsoukas, Christos, et al.
Published: (2023)
by: Matsoukas, Christos, et al.
Published: (2023)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
by: Han, Donghoon, et al.
Published: (2023)
by: Han, Donghoon, et al.
Published: (2023)
Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay
by: Tong, Jin, et al.
Published: (2026)
by: Tong, Jin, et al.
Published: (2026)
Token Cropr: Faster ViTs for Quite a Few Tasks
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
by: Hwang, Dongyoon, et al.
Published: (2024)
by: Hwang, Dongyoon, et al.
Published: (2024)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
by: Siméoni, Oriane, et al.
Published: (2023)
by: Siméoni, Oriane, et al.
Published: (2023)
VIVID-Med: LLM-Supervised Structured Pretraining for Deployable Medical ViTs
by: Wang, Xiyao, et al.
Published: (2026)
by: Wang, Xiyao, et al.
Published: (2026)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
by: Alvetreti, Federico, et al.
Published: (2025)
by: Alvetreti, Federico, et al.
Published: (2025)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP
by: Balasubramanian, Sriram, et al.
Published: (2024)
by: Balasubramanian, Sriram, et al.
Published: (2024)
Parameter Efficient Fine-tuning of Self-supervised ViTs without Catastrophic Forgetting
by: Bafghi, Reza Akbarian, et al.
Published: (2024)
by: Bafghi, Reza Akbarian, et al.
Published: (2024)
ViTs: Teaching Machines to See Time Series Anomalies Like Human Experts
by: Wang, Zexin, et al.
Published: (2025)
by: Wang, Zexin, et al.
Published: (2025)
Causality $\neq$ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
by: Huang, Lianghuan, et al.
Published: (2025)
by: Huang, Lianghuan, et al.
Published: (2025)
ViT Registers and Fractal ViT
by: Chou, Jason Chuan-Chih, et al.
Published: (2026)
by: Chou, Jason Chuan-Chih, et al.
Published: (2026)
Similar Items
-
SaRPFF: A Self-Attention with Register-based Pyramid Feature Fusion module for enhanced RLD detection
by: Haruna, Yunusa, et al.
Published: (2024) -
KonvLiNA: Integrating Kolmogorov-Arnold Network with Linear Nyström Attention for feature fusion in Crop Field Detection
by: Yunusa, Haruna, et al.
Published: (2024) -
iiANET: Inception Inspired Attention Hybrid Network for efficient Long-Range Dependency
by: Yunusa, Haruna, et al.
Published: (2024) -
VulCatch: Enhancing Binary Vulnerability Detection through CodeT5 Decompilation and KAN Advanced Feature Extraction
by: Chukkol, Abdulrahman Hamman Adama, et al.
Published: (2024) -
GateMABSA: Aspect-Image Gated Fusion for Multimodal Aspect-based Sentiment Analysis
by: Lawan, Adamu, et al.
Published: (2025)