Combined CNN and ViT features off-the-shelf: Another astounding baseline for recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Alonso-Fernandez, Fernando, Hernandez-Diaz, Kevin, Tiwari, Prayag, Bigun, Josef |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Deep Network Pruning: A Comparative Study on CNNs in Face Recognition
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2024)
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2024)
Understanding and Improving CNNs with Complex Structure Tensor: A Biometrics Study
por: Hernandez-Diaz, Kevin, et al.
Publicado: (2024)
por: Hernandez-Diaz, Kevin, et al.
Publicado: (2024)
SqueezeFacePoseNet: Lightweight Face Verification Across Different Poses for Mobile Platforms
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2020)
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2020)
Total Least Square Optimal Analytic Signal by Structure Tensor for N-D images
por: Bigun, Josef, et al.
Publicado: (2020)
por: Bigun, Josef, et al.
Publicado: (2020)
Leveraging Large-Scale Face Datasets for Deep Periocular Recognition via Ocular Cropping
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2025)
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2025)
Note on the Construction of Structure Tensor
por: Bigun, Josef, et al.
Publicado: (2025)
por: Bigun, Josef, et al.
Publicado: (2025)
Exploring Complementarity and Explainability in CNNs for Periocular Verification Across Acquisition Distances
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2025)
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2025)
RepViT: Revisiting Mobile CNN From ViT Perspective
por: Wang, Ao, et al.
Publicado: (2023)
por: Wang, Ao, et al.
Publicado: (2023)
Overtake Detection in Trucks Using CAN Bus Signals: A Comparative Study of Machine Learning Methods
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2025)
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2025)
Exploring the correlation between the type of music and the emotions evoked: A study using subjective questionnaires and EEG
por: Jankowska, Jelizaveta, et al.
Publicado: (2025)
por: Jankowska, Jelizaveta, et al.
Publicado: (2025)
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
por: Amangeldi, Aidar, et al.
Publicado: (2025)
por: Amangeldi, Aidar, et al.
Publicado: (2025)
OmniPatch: A Universal Adversarial Patch for ViT-CNN Cross-Architecture Transfer in Semantic Segmentation
por: Aggarwal, Aarush, et al.
Publicado: (2026)
por: Aggarwal, Aarush, et al.
Publicado: (2026)
Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection
por: Gao, Xiangyu, et al.
Publicado: (2025)
por: Gao, Xiangyu, et al.
Publicado: (2025)
Hybrid CNN-ViT Framework for Motion-Blurred Scene Text Restoration
por: Rashid, Umar, et al.
Publicado: (2025)
por: Rashid, Umar, et al.
Publicado: (2025)
Learning CNN on ViT: A Hybrid Model to Explicitly Class-specific Boundaries for Domain Adaptation
por: Ngo, Ba Hung, et al.
Publicado: (2024)
por: Ngo, Ba Hung, et al.
Publicado: (2024)
A Hybrid Framework Bridging CNN and ViT based on Theory of Evidence for Diabetic Retinopathy Grading
por: Qiu, Junlai, et al.
Publicado: (2025)
por: Qiu, Junlai, et al.
Publicado: (2025)
Deeper Inside Deep ViT
por: Hong, Sungrae
Publicado: (2025)
por: Hong, Sungrae
Publicado: (2025)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
por: Zhong, Yunshan, et al.
Publicado: (2023)
por: Zhong, Yunshan, et al.
Publicado: (2023)
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
por: Hernandez, Juan Manuel, et al.
Publicado: (2026)
por: Hernandez, Juan Manuel, et al.
Publicado: (2026)
Pretrained ViTs Yield Versatile Representations For Medical Images
por: Matsoukas, Christos, et al.
Publicado: (2023)
por: Matsoukas, Christos, et al.
Publicado: (2023)
A Hybrid CNN-ViT-GNN Framework with GAN-Based Augmentation for Intelligent Weed Detection in Precision Agriculture
por: V, Pandiyaraju, et al.
Publicado: (2025)
por: V, Pandiyaraju, et al.
Publicado: (2025)
Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentation
por: Tang, Fenghe, et al.
Publicado: (2025)
por: Tang, Fenghe, et al.
Publicado: (2025)
Rethinking Random Masking in Self-Distillation on ViT
por: Seong, Jihyeon, et al.
Publicado: (2025)
por: Seong, Jihyeon, et al.
Publicado: (2025)
YOLO-Former: YOLO Shakes Hand With ViT
por: Khoramdel, Javad, et al.
Publicado: (2024)
por: Khoramdel, Javad, et al.
Publicado: (2024)
Your ViT is Secretly an Image Segmentation Model
por: Kerssies, Tommie, et al.
Publicado: (2025)
por: Kerssies, Tommie, et al.
Publicado: (2025)
ViT-5: Vision Transformers for The Mid-2020s
por: Wang, Feng, et al.
Publicado: (2026)
por: Wang, Feng, et al.
Publicado: (2026)
CNN-ViT Hybrid for Pneumonia Detection: Theory and Empiric on Limited Data without Pretraining
por: Basnet, Prashant Singh, et al.
Publicado: (2025)
por: Basnet, Prashant Singh, et al.
Publicado: (2025)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
por: Chattopadhyay, Nandish, et al.
Publicado: (2026)
por: Chattopadhyay, Nandish, et al.
Publicado: (2026)
PaW-ViT: A Patch-based Warping Vision Transformer for Robust Ear Verification
por: Arun, Deeksha, et al.
Publicado: (2026)
por: Arun, Deeksha, et al.
Publicado: (2026)
ViTCAE: ViT-based Class-conditioned Autoencoder
por: Jebraeeli, Vahid, et al.
Publicado: (2025)
por: Jebraeeli, Vahid, et al.
Publicado: (2025)
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
por: Puy, Gilles, et al.
Publicado: (2026)
por: Puy, Gilles, et al.
Publicado: (2026)
Applying ViT in Generalized Few-shot Semantic Segmentation
por: Geng, Liyuan, et al.
Publicado: (2024)
por: Geng, Liyuan, et al.
Publicado: (2024)
One-Shot Multilingual Font Generation Via ViT
por: Wang, Zhiheng, et al.
Publicado: (2024)
por: Wang, Zhiheng, et al.
Publicado: (2024)
ViT$^3$: Unlocking Test-Time Training in Vision
por: Han, Dongchen, et al.
Publicado: (2025)
por: Han, Dongchen, et al.
Publicado: (2025)
ACC-ViT : Atrous Convolution's Comeback in Vision Transformers
por: Ibtehaz, Nabil, et al.
Publicado: (2024)
por: Ibtehaz, Nabil, et al.
Publicado: (2024)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
por: Zhu, Chen, et al.
Publicado: (2025)
por: Zhu, Chen, et al.
Publicado: (2025)
Register and [CLS] tokens yield a decoupling of local and global features in large ViTs
por: Lappe, Alexander, et al.
Publicado: (2025)
por: Lappe, Alexander, et al.
Publicado: (2025)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
por: Shah, Arya, et al.
Publicado: (2025)
por: Shah, Arya, et al.
Publicado: (2025)
H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction
por: Li, Xueyang, et al.
Publicado: (2025)
por: Li, Xueyang, et al.
Publicado: (2025)
U-REPA: Aligning Diffusion U-Nets to ViTs
por: Tian, Yuchuan, et al.
Publicado: (2025)
por: Tian, Yuchuan, et al.
Publicado: (2025)
Ejemplares similares
-
Deep Network Pruning: A Comparative Study on CNNs in Face Recognition
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2024) -
Understanding and Improving CNNs with Complex Structure Tensor: A Biometrics Study
por: Hernandez-Diaz, Kevin, et al.
Publicado: (2024) -
SqueezeFacePoseNet: Lightweight Face Verification Across Different Poses for Mobile Platforms
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2020) -
Total Least Square Optimal Analytic Signal by Structure Tensor for N-D images
por: Bigun, Josef, et al.
Publicado: (2020) -
Leveraging Large-Scale Face Datasets for Deep Periocular Recognition via Ocular Cropping
por: Alonso-Fernandez, Fernando, et al.
Publicado: (2025)