MSCViT: A Small-size ViT architecture with Multi-Scale Self-Attention Mechanism for Tiny Datasets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Bowei, Zhang, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
von: Amangeldi, Aidar, et al.
Veröffentlicht: (2025)
von: Amangeldi, Aidar, et al.
Veröffentlicht: (2025)
Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
von: You, Haoran, et al.
Veröffentlicht: (2022)
von: You, Haoran, et al.
Veröffentlicht: (2022)
Rethinking Random Masking in Self-Distillation on ViT
von: Seong, Jihyeon, et al.
Veröffentlicht: (2025)
von: Seong, Jihyeon, et al.
Veröffentlicht: (2025)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
DeS3: Adaptive Attention-driven Self and Soft Shadow Removal using ViT Similarity
von: Jin, Yeying, et al.
Veröffentlicht: (2022)
von: Jin, Yeying, et al.
Veröffentlicht: (2022)
ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers
von: Sinha, Sanchit, et al.
Veröffentlicht: (2025)
von: Sinha, Sanchit, et al.
Veröffentlicht: (2025)
LAMM-ViT: AI Face Detection via Layer-Aware Modulation of Region-Guided Attention
von: Zhang, Jiangling, et al.
Veröffentlicht: (2025)
von: Zhang, Jiangling, et al.
Veröffentlicht: (2025)
Tiny-ViT: A Compact Vision Transformer for Efficient and Explainable Potato Leaf Disease Classification
von: Mia, Shakil, et al.
Veröffentlicht: (2026)
von: Mia, Shakil, et al.
Veröffentlicht: (2026)
Alias-Free ViT: Fractional Shift Invariance via Linear Attention
von: Michaeli, Hagay, et al.
Veröffentlicht: (2025)
von: Michaeli, Hagay, et al.
Veröffentlicht: (2025)
Deeper Inside Deep ViT
von: Hong, Sungrae
Veröffentlicht: (2025)
von: Hong, Sungrae
Veröffentlicht: (2025)
Exploring Plain ViT Reconstruction for Multi-class Unsupervised Anomaly Detection
von: Zhang, Jiangning, et al.
Veröffentlicht: (2023)
von: Zhang, Jiangning, et al.
Veröffentlicht: (2023)
ViT-5: Vision Transformers for The Mid-2020s
von: Wang, Feng, et al.
Veröffentlicht: (2026)
von: Wang, Feng, et al.
Veröffentlicht: (2026)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
von: Zhong, Yunshan, et al.
Veröffentlicht: (2023)
von: Zhong, Yunshan, et al.
Veröffentlicht: (2023)
MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
AFIDAF: Alternating Fourier and Image Domain Adaptive Filters as an Efficient Alternative to Attention in ViTs
von: Zheng, Yunling, et al.
Veröffentlicht: (2024)
von: Zheng, Yunling, et al.
Veröffentlicht: (2024)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
von: Shah, Arya, et al.
Veröffentlicht: (2025)
von: Shah, Arya, et al.
Veröffentlicht: (2025)
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
von: Zhang, Tianfang, et al.
Veröffentlicht: (2024)
von: Zhang, Tianfang, et al.
Veröffentlicht: (2024)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
von: Siméoni, Oriane, et al.
Veröffentlicht: (2023)
von: Siméoni, Oriane, et al.
Veröffentlicht: (2023)
ViT-Lens: Towards Omni-modal Representations
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
von: Yao, Ting, et al.
Veröffentlicht: (2024)
von: Yao, Ting, et al.
Veröffentlicht: (2024)
RepViT: Revisiting Mobile CNN From ViT Perspective
von: Wang, Ao, et al.
Veröffentlicht: (2023)
von: Wang, Ao, et al.
Veröffentlicht: (2023)
H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction
von: Li, Xueyang, et al.
Veröffentlicht: (2025)
von: Li, Xueyang, et al.
Veröffentlicht: (2025)
ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
Retinal Malady Classification using AI: A novel ViT-SVM combination architecture
von: Jha, Shashwat, et al.
Veröffentlicht: (2026)
von: Jha, Shashwat, et al.
Veröffentlicht: (2026)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
MedSAM-CA: A CNN-Augmented ViT with Attention-Enhanced Multi-Scale Fusion for Medical Image Segmentation
von: Tian, Peiting, et al.
Veröffentlicht: (2025)
von: Tian, Peiting, et al.
Veröffentlicht: (2025)
A Lightweight Convolution and Vision Transformer integrated model with Multi-scale Self-attention Mechanism
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
Make Your ViT-based Multi-view 3D Detectors Faster via Token Compression
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2024)
DC-ViT: Modulating Spatial and Channel Interactions for Multi-Channel Images
von: Marikkar, Umar, et al.
Veröffentlicht: (2026)
von: Marikkar, Umar, et al.
Veröffentlicht: (2026)
Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors
von: Ji, Mingqian, et al.
Veröffentlicht: (2026)
von: Ji, Mingqian, et al.
Veröffentlicht: (2026)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025)
Poseidon: A ViT-based Architecture for Multi-Frame Pose Estimation with Adaptive Frame Weighting and Multi-Scale Feature Fusion
von: Pace, Cesare Davide, et al.
Veröffentlicht: (2025)
von: Pace, Cesare Davide, et al.
Veröffentlicht: (2025)
Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability
von: Liu, Jiani, et al.
Veröffentlicht: (2025)
von: Liu, Jiani, et al.
Veröffentlicht: (2025)
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
von: Qiu, Congpei, et al.
Veröffentlicht: (2026)
von: Qiu, Congpei, et al.
Veröffentlicht: (2026)
Your ViT is Secretly an Image Segmentation Model
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
YOLO-Former: YOLO Shakes Hand With ViT
von: Khoramdel, Javad, et al.
Veröffentlicht: (2024)
von: Khoramdel, Javad, et al.
Veröffentlicht: (2024)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
TinyViM: Frequency Decoupling for Tiny Hybrid Vision Mamba
von: Ma, Xiaowen, et al.
Veröffentlicht: (2024)
von: Ma, Xiaowen, et al.
Veröffentlicht: (2024)
Parameter Efficient Fine-tuning of Self-supervised ViTs without Catastrophic Forgetting
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2024)
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2024)
ViTCAE: ViT-based Class-conditioned Autoencoder
von: Jebraeeli, Vahid, et al.
Veröffentlicht: (2025)
von: Jebraeeli, Vahid, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
von: Amangeldi, Aidar, et al.
Veröffentlicht: (2025) -
Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
von: You, Haoran, et al.
Veröffentlicht: (2022) -
Rethinking Random Masking in Self-Distillation on ViT
von: Seong, Jihyeon, et al.
Veröffentlicht: (2025) -
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
von: Li, Yifan, et al.
Veröffentlicht: (2026) -
DeS3: Adaptive Attention-driven Self and Soft Shadow Removal using ViT Similarity
von: Jin, Yeying, et al.
Veröffentlicht: (2022)