Do All Vision Transformers Need Registers? A Cross-Architectural Reassessment
Fuente:
arXiv
Salvato in:
| Autori principali: | Baxevanakis, Spiros, Karageorgis, Platon, Dravilas, Ioannis, Szewczyk, Konrad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
Visible Iris Area as a Quality Metric for Reliable Iris Recognition Under Pupil Dilation and Eyelid Occlusion
di: Pessaud, Jack, et al.
Pubblicazione: (2025)
di: Pessaud, Jack, et al.
Pubblicazione: (2025)
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
di: Deka, Dawar Jyoti, et al.
Pubblicazione: (2026)
di: Deka, Dawar Jyoti, et al.
Pubblicazione: (2026)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
di: Wang, Yiming, et al.
Pubblicazione: (2026)
di: Wang, Yiming, et al.
Pubblicazione: (2026)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
di: Li, Yayuan, et al.
Pubblicazione: (2025)
di: Li, Yayuan, et al.
Pubblicazione: (2025)
SelvaBox: A high-resolution dataset for tropical tree crown detection
di: Baudchon, Hugo, et al.
Pubblicazione: (2025)
di: Baudchon, Hugo, et al.
Pubblicazione: (2025)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
di: Deng, Pei, et al.
Pubblicazione: (2025)
di: Deng, Pei, et al.
Pubblicazione: (2025)
Image-Based Leopard Seal Recognition: Approaches and Challenges in Current Automated Systems
di: Salazar, Jorge Yero, et al.
Pubblicazione: (2024)
di: Salazar, Jorge Yero, et al.
Pubblicazione: (2024)
Dense Motion Captioning
di: Xu, Shiyao, et al.
Pubblicazione: (2025)
di: Xu, Shiyao, et al.
Pubblicazione: (2025)
TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
di: Lee, Byung Hoon, et al.
Pubblicazione: (2025)
di: Lee, Byung Hoon, et al.
Pubblicazione: (2025)
CoMatcher: Multi-View Collaborative Feature Matching
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
di: Zhang, Jintao, et al.
Pubblicazione: (2025)
MoDE: Mixture of Diffusion Experts for Any Occluded Face Recognition
di: Fan, Qiannan, et al.
Pubblicazione: (2025)
di: Fan, Qiannan, et al.
Pubblicazione: (2025)
NeuroGaze-Distill: Brain-informed Distillation and Depression-Inspired Geometric Priors for Robust Facial Emotion Recognition
di: Li, Zilin, et al.
Pubblicazione: (2025)
di: Li, Zilin, et al.
Pubblicazione: (2025)
PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
di: Yang, Zongyou, et al.
Pubblicazione: (2025)
di: Yang, Zongyou, et al.
Pubblicazione: (2025)
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
di: Diller, Christian, et al.
Pubblicazione: (2023)
di: Diller, Christian, et al.
Pubblicazione: (2023)
SelvaMask: Segmenting Trees in Tropical Forests and Beyond
di: Duguay, Simon-Olivier, et al.
Pubblicazione: (2026)
di: Duguay, Simon-Olivier, et al.
Pubblicazione: (2026)
Dynamic Arthroscopic Navigation System for Anterior Cruciate Ligament Reconstruction Based on Multi-level Memory Architecture
di: Wang, Shuo, et al.
Pubblicazione: (2025)
di: Wang, Shuo, et al.
Pubblicazione: (2025)
Distributed Intelligent System Architecture for UAV-Assisted Monitoring of Wind Energy Infrastructure
di: Svystun, Serhii, et al.
Pubblicazione: (2024)
di: Svystun, Serhii, et al.
Pubblicazione: (2024)
Capacity Constraint Analysis Using Object Detection for Smart Manufacturing
di: Ahmad, Hafiz Mughees, et al.
Pubblicazione: (2024)
di: Ahmad, Hafiz Mughees, et al.
Pubblicazione: (2024)
SH17: A Dataset for Human Safety and Personal Protective Equipment Detection in Manufacturing Industry
di: Ahmad, Hafiz Mughees, et al.
Pubblicazione: (2024)
di: Ahmad, Hafiz Mughees, et al.
Pubblicazione: (2024)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
di: Huang, Yian, et al.
Pubblicazione: (2026)
di: Huang, Yian, et al.
Pubblicazione: (2026)
Habitat Classification from Ground-Level Imagery Using Deep Neural Networks
di: Shi, Hongrui, et al.
Pubblicazione: (2025)
di: Shi, Hongrui, et al.
Pubblicazione: (2025)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
di: Brovko, D. V.
Pubblicazione: (2025)
di: Brovko, D. V.
Pubblicazione: (2025)
Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
di: Dong, Haohua, et al.
Pubblicazione: (2025)
di: Dong, Haohua, et al.
Pubblicazione: (2025)
HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding
di: Jin, Haopeng, et al.
Pubblicazione: (2026)
di: Jin, Haopeng, et al.
Pubblicazione: (2026)
Towards Accurate and Efficient Waste Image Classification: A Hybrid Deep Learning and Machine Learning Approach
di: Nguyen, Ngoc-Bao-Quang, et al.
Pubblicazione: (2025)
di: Nguyen, Ngoc-Bao-Quang, et al.
Pubblicazione: (2025)
Efficient Temporally-Aware DeepFake Detection using H.264 Motion Vectors
di: Grönquist, Peter, et al.
Pubblicazione: (2023)
di: Grönquist, Peter, et al.
Pubblicazione: (2023)
FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video Observations
di: Diller, Christian, et al.
Pubblicazione: (2022)
di: Diller, Christian, et al.
Pubblicazione: (2022)
Detecting AI-Generated Videos with Spiking Neural Networks
di: Jang, Minsuk, et al.
Pubblicazione: (2026)
di: Jang, Minsuk, et al.
Pubblicazione: (2026)
EgoTraj: Real-World Egocentric Human Trajectory Dataset for Multimodal Prediction
di: Yehia, Ahmad, et al.
Pubblicazione: (2026)
di: Yehia, Ahmad, et al.
Pubblicazione: (2026)
Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
Distance Estimation in Outdoor Driving Environments Using Phase-only Correlation Method with Event Cameras
di: Kobayashi, Masataka, et al.
Pubblicazione: (2025)
di: Kobayashi, Masataka, et al.
Pubblicazione: (2025)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
di: Zinnen, Mathias, et al.
Pubblicazione: (2025)
di: Zinnen, Mathias, et al.
Pubblicazione: (2025)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
di: Kalkanli, Beyza, et al.
Pubblicazione: (2026)
di: Kalkanli, Beyza, et al.
Pubblicazione: (2026)
Transfer-learning for video classification: Video Swin Transformer on multiple domains
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2022)
di: Oliveira, Daniel A. P., et al.
Pubblicazione: (2022)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
di: Gupta, Sunny, et al.
Pubblicazione: (2024)
di: Gupta, Sunny, et al.
Pubblicazione: (2024)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
License Plate Detection and Character Recognition Using Deep Learning and Font Evaluation
di: Vargoorani, Zahra Ebrahimi, et al.
Pubblicazione: (2024)
di: Vargoorani, Zahra Ebrahimi, et al.
Pubblicazione: (2024)
LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar Animation
di: Wei, Hualiang, et al.
Pubblicazione: (2026)
di: Wei, Hualiang, et al.
Pubblicazione: (2026)
How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?
di: Costa, Daniel da Silva, et al.
Pubblicazione: (2026)
di: Costa, Daniel da Silva, et al.
Pubblicazione: (2026)
Documenti analoghi
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025) -
Visible Iris Area as a Quality Metric for Reliable Iris Recognition Under Pupil Dilation and Eyelid Occlusion
di: Pessaud, Jack, et al.
Pubblicazione: (2025) -
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
di: Deka, Dawar Jyoti, et al.
Pubblicazione: (2026) -
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
di: Wang, Yiming, et al.
Pubblicazione: (2026) -
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
di: Li, Yayuan, et al.
Pubblicazione: (2025)