RepViT: Revisiting Mobile CNN From ViT Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Ao, Chen, Hui, Lin, Zijia, Han, Jungong, Ding, Guiguang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RepViT-SAM: Towards Real-Time Segmenting Anything
von: Wang, Ao, et al.
Veröffentlicht: (2023)
von: Wang, Ao, et al.
Veröffentlicht: (2023)
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs
von: Wang, Ao, et al.
Veröffentlicht: (2023)
von: Wang, Ao, et al.
Veröffentlicht: (2023)
LSNet: See Large, Focus Small
von: Wang, Ao, et al.
Veröffentlicht: (2025)
von: Wang, Ao, et al.
Veröffentlicht: (2025)
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
von: Wang, Ao, et al.
Veröffentlicht: (2024)
von: Wang, Ao, et al.
Veröffentlicht: (2024)
YOLOE: Real-Time Seeing Anything
von: Wang, Ao, et al.
Veröffentlicht: (2025)
von: Wang, Ao, et al.
Veröffentlicht: (2025)
YOLOv10: Real-Time End-to-End Object Detection
von: Wang, Ao, et al.
Veröffentlicht: (2024)
von: Wang, Ao, et al.
Veröffentlicht: (2024)
Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentation
von: Tang, Fenghe, et al.
Veröffentlicht: (2025)
von: Tang, Fenghe, et al.
Veröffentlicht: (2025)
PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
von: Wang, Ao, et al.
Veröffentlicht: (2024)
von: Wang, Ao, et al.
Veröffentlicht: (2024)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
von: Zhong, Yunshan, et al.
Veröffentlicht: (2023)
von: Zhong, Yunshan, et al.
Veröffentlicht: (2023)
Context Enhancement with Reconstruction as Sequence for Unified Unsupervised Anomaly Detection
von: Yang, Hui-Yue, et al.
Veröffentlicht: (2024)
von: Yang, Hui-Yue, et al.
Veröffentlicht: (2024)
Rethinking Random Masking in Self-Distillation on ViT
von: Seong, Jihyeon, et al.
Veröffentlicht: (2025)
von: Seong, Jihyeon, et al.
Veröffentlicht: (2025)
Deeper Inside Deep ViT
von: Hong, Sungrae
Veröffentlicht: (2025)
von: Hong, Sungrae
Veröffentlicht: (2025)
Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning
von: Yang, Hui-Yue, et al.
Veröffentlicht: (2024)
von: Yang, Hui-Yue, et al.
Veröffentlicht: (2024)
Combined CNN and ViT features off-the-shelf: Another astounding baseline for recognition
von: Alonso-Fernandez, Fernando, et al.
Veröffentlicht: (2024)
von: Alonso-Fernandez, Fernando, et al.
Veröffentlicht: (2024)
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
von: Amangeldi, Aidar, et al.
Veröffentlicht: (2025)
von: Amangeldi, Aidar, et al.
Veröffentlicht: (2025)
YOLO-UniOW: Efficient Universal Open-World Object Detection
von: Liu, Lihao, et al.
Veröffentlicht: (2024)
von: Liu, Lihao, et al.
Veröffentlicht: (2024)
Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection
von: Gao, Xiangyu, et al.
Veröffentlicht: (2025)
von: Gao, Xiangyu, et al.
Veröffentlicht: (2025)
RepViT-CXR: A Channel Replication Strategy for Vision Transformers in Chest X-ray Tuberculosis and Pneumonia Classification
von: Ahmed, Faisal
Veröffentlicht: (2025)
von: Ahmed, Faisal
Veröffentlicht: (2025)
ViTCAE: ViT-based Class-conditioned Autoencoder
von: Jebraeeli, Vahid, et al.
Veröffentlicht: (2025)
von: Jebraeeli, Vahid, et al.
Veröffentlicht: (2025)
Hybrid CNN-ViT Framework for Motion-Blurred Scene Text Restoration
von: Rashid, Umar, et al.
Veröffentlicht: (2025)
von: Rashid, Umar, et al.
Veröffentlicht: (2025)
A Hybrid Framework Bridging CNN and ViT based on Theory of Evidence for Diabetic Retinopathy Grading
von: Qiu, Junlai, et al.
Veröffentlicht: (2025)
von: Qiu, Junlai, et al.
Veröffentlicht: (2025)
ViT-1.58b: Mobile Vision Transformers in the 1-bit Era
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2024)
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2024)
ViT$^3$: Unlocking Test-Time Training in Vision
von: Han, Dongchen, et al.
Veröffentlicht: (2025)
von: Han, Dongchen, et al.
Veröffentlicht: (2025)
ViT-5: Vision Transformers for The Mid-2020s
von: Wang, Feng, et al.
Veröffentlicht: (2026)
von: Wang, Feng, et al.
Veröffentlicht: (2026)
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
PYRA: Parallel Yielding Re-Activation for Training-Inference Efficient Task Adaptation
von: Xiong, Yizhe, et al.
Veröffentlicht: (2024)
von: Xiong, Yizhe, et al.
Veröffentlicht: (2024)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
von: Tonmoy, Moshiur Rahman, et al.
Veröffentlicht: (2025)
von: Tonmoy, Moshiur Rahman, et al.
Veröffentlicht: (2025)
Learn from the Learnt: Source-Free Active Domain Adaptation via Contrastive Sampling and Visual Persistence
von: Lyu, Mengyao, et al.
Veröffentlicht: (2024)
von: Lyu, Mengyao, et al.
Veröffentlicht: (2024)
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
von: Zhang, Tianfang, et al.
Veröffentlicht: (2024)
von: Zhang, Tianfang, et al.
Veröffentlicht: (2024)
Learning CNN on ViT: A Hybrid Model to Explicitly Class-specific Boundaries for Domain Adaptation
von: Ngo, Ba Hung, et al.
Veröffentlicht: (2024)
von: Ngo, Ba Hung, et al.
Veröffentlicht: (2024)
One-Shot Multilingual Font Generation Via ViT
von: Wang, Zhiheng, et al.
Veröffentlicht: (2024)
von: Wang, Zhiheng, et al.
Veröffentlicht: (2024)
YOLO-Former: YOLO Shakes Hand With ViT
von: Khoramdel, Javad, et al.
Veröffentlicht: (2024)
von: Khoramdel, Javad, et al.
Veröffentlicht: (2024)
Your ViT is Secretly an Image Segmentation Model
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation
von: Xiong, Yizhe, et al.
Veröffentlicht: (2025)
von: Xiong, Yizhe, et al.
Veröffentlicht: (2025)
HydraViT: Stacking Heads for a Scalable ViT
von: Haberer, Janek, et al.
Veröffentlicht: (2024)
von: Haberer, Janek, et al.
Veröffentlicht: (2024)
U-REPA: Aligning Diffusion U-Nets to ViTs
von: Tian, Yuchuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuchuan, et al.
Veröffentlicht: (2025)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
von: Chen, Lu, et al.
Veröffentlicht: (2025)
von: Chen, Lu, et al.
Veröffentlicht: (2025)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
von: Shah, Arya, et al.
Veröffentlicht: (2025)
von: Shah, Arya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RepViT-SAM: Towards Real-Time Segmenting Anything
von: Wang, Ao, et al.
Veröffentlicht: (2023) -
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs
von: Wang, Ao, et al.
Veröffentlicht: (2023) -
LSNet: See Large, Focus Small
von: Wang, Ao, et al.
Veröffentlicht: (2025) -
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
von: Wang, Ao, et al.
Veröffentlicht: (2024) -
YOLOE: Real-Time Seeing Anything
von: Wang, Ao, et al.
Veröffentlicht: (2025)