Revisiting [CLS] and Patch Token Interaction in Vision Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Marouani, Alexis, Siméoni, Oriane, Jégou, Hervé, Bojanowski, Piotr, Vo, Huy V. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unveiling Text in Challenging Stone Inscriptions: A Character-Context-Aware Patching Strategy for Binarization
di: Jena, Pratyush, et al.
Pubblicazione: (2026)
di: Jena, Pratyush, et al.
Pubblicazione: (2026)
LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
di: Jia, Yongju, et al.
Pubblicazione: (2025)
di: Jia, Yongju, et al.
Pubblicazione: (2025)
A Spitting Image: Modular Superpixel Tokenization in Vision Transformers
di: Aasan, Marius, et al.
Pubblicazione: (2024)
di: Aasan, Marius, et al.
Pubblicazione: (2024)
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
di: Yasuno, Takato
Pubblicazione: (2026)
di: Yasuno, Takato
Pubblicazione: (2026)
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
di: Chaybouti, Sofian, et al.
Pubblicazione: (2025)
di: Chaybouti, Sofian, et al.
Pubblicazione: (2025)
Learning Unified Representation of 3D Gaussian Splatting
di: Xin, Yuelin, et al.
Pubblicazione: (2025)
di: Xin, Yuelin, et al.
Pubblicazione: (2025)
VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
di: Chen, Zhipeng, et al.
Pubblicazione: (2024)
di: Chen, Zhipeng, et al.
Pubblicazione: (2024)
HySparK: Hybrid Sparse Masking for Large Scale Medical Image Pre-Training
di: Tang, Fenghe, et al.
Pubblicazione: (2024)
di: Tang, Fenghe, et al.
Pubblicazione: (2024)
Mobile-Ready Automated Triage of Diabetic Retinopathy Using Digital Fundus Images
di: Joshi, Aadi, et al.
Pubblicazione: (2026)
di: Joshi, Aadi, et al.
Pubblicazione: (2026)
Differentiable Hierarchical Visual Tokenization
di: Aasan, Marius, et al.
Pubblicazione: (2025)
di: Aasan, Marius, et al.
Pubblicazione: (2025)
Frequency-Decomposed INR for NIR-Assisted Low-Light RGB Image Denoising
di: Shi, Ligen, et al.
Pubblicazione: (2026)
di: Shi, Ligen, et al.
Pubblicazione: (2026)
Neural Fields for 3D Tracking of Anatomy and Surgical Instruments in Monocular Laparoscopic Video Clips
di: Gerats, Beerend G. A., et al.
Pubblicazione: (2024)
di: Gerats, Beerend G. A., et al.
Pubblicazione: (2024)
A Hierarchical Self-Consistent Regularization Approach to Satellite Image Time Series Classification
di: Weikmann, Giulio, et al.
Pubblicazione: (2025)
di: Weikmann, Giulio, et al.
Pubblicazione: (2025)
Learning to Expand Images for Efficient Visual Autoregressive Modeling
di: Yang, Ruiqing, et al.
Pubblicazione: (2025)
di: Yang, Ruiqing, et al.
Pubblicazione: (2025)
Scalable Face Security Vision Foundation Model for Deepfake, Diffusion, and Spoofing Detection
di: Wang, Gaojian, et al.
Pubblicazione: (2025)
di: Wang, Gaojian, et al.
Pubblicazione: (2025)
Cora: Correspondence-aware image editing using few step diffusion
di: Alimohammadi, Amirhossein, et al.
Pubblicazione: (2025)
di: Alimohammadi, Amirhossein, et al.
Pubblicazione: (2025)
Pointing-Based Object Recognition
di: Hajdúch, Lukáš, et al.
Pubblicazione: (2026)
di: Hajdúch, Lukáš, et al.
Pubblicazione: (2026)
Prototype-Guided Concept Erasure in Diffusion Models
di: Cai, Yuze, et al.
Pubblicazione: (2026)
di: Cai, Yuze, et al.
Pubblicazione: (2026)
Supervised Contrastive Learning for Few-Shot AI-Generated Image Detection and Attribution
di: Urueña, Jaime Álvarez, et al.
Pubblicazione: (2025)
di: Urueña, Jaime Álvarez, et al.
Pubblicazione: (2025)
Label Delay in Online Continual Learning
di: Csaba, Botos, et al.
Pubblicazione: (2023)
di: Csaba, Botos, et al.
Pubblicazione: (2023)
Neural Implicit Morphing of Face Images
di: Schardong, Guilherme, et al.
Pubblicazione: (2023)
di: Schardong, Guilherme, et al.
Pubblicazione: (2023)
RealHD: A High-Quality Dataset for Robust Detection of State-of-the-Art AI-Generated Images
di: Yu, Hanzhe, et al.
Pubblicazione: (2026)
di: Yu, Hanzhe, et al.
Pubblicazione: (2026)
A Novel Global Context-aware Deep Neural Network for Enhanced Brain Tumor Segmentation using Magnetic Resonance Images
di: Mukherjee, Sourjya, et al.
Pubblicazione: (2026)
di: Mukherjee, Sourjya, et al.
Pubblicazione: (2026)
When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation
di: Frochte, Jörg
Pubblicazione: (2026)
di: Frochte, Jörg
Pubblicazione: (2026)
FAME: Feature Activation Map Explanation on Image Classification and Face Recognition
di: Zhang, Xinyi, et al.
Pubblicazione: (2026)
di: Zhang, Xinyi, et al.
Pubblicazione: (2026)
UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis Augmentation
di: Zhao, Zhihao, et al.
Pubblicazione: (2025)
di: Zhao, Zhihao, et al.
Pubblicazione: (2025)
GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations
di: Liu, Zeping, et al.
Pubblicazione: (2025)
di: Liu, Zeping, et al.
Pubblicazione: (2025)
MienCap: Realtime Performance-Based Facial Animation with Live Mood Dynamics
di: Pan, Ye, et al.
Pubblicazione: (2025)
di: Pan, Ye, et al.
Pubblicazione: (2025)
Generating real-time detailed ground visualisations from sparse aerial point clouds
di: Murray, Aidan, et al.
Pubblicazione: (2025)
di: Murray, Aidan, et al.
Pubblicazione: (2025)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
di: Li, Jinhao, et al.
Pubblicazione: (2024)
di: Li, Jinhao, et al.
Pubblicazione: (2024)
Symmetry Awareness Encoded Deep Learning Framework for Brain Imaging Analysis
di: Ma, Yang, et al.
Pubblicazione: (2024)
di: Ma, Yang, et al.
Pubblicazione: (2024)
STimage-1K4M: A histopathology image-gene expression dataset for spatial transcriptomics
di: Chen, Jiawen, et al.
Pubblicazione: (2024)
di: Chen, Jiawen, et al.
Pubblicazione: (2024)
ViFiCon: Vision and Wireless Association Via Self-Supervised Contrastive Learning
di: Meegan, Nicholas, et al.
Pubblicazione: (2022)
di: Meegan, Nicholas, et al.
Pubblicazione: (2022)
Implicitly Learned Neural Phase Functions for Basis-Free Point Spread Function Engineering
di: Valouev, Aleksey
Pubblicazione: (2024)
di: Valouev, Aleksey
Pubblicazione: (2024)
Meta Co-Training: Two Views are Better than One
di: Rothenberger, Jay C., et al.
Pubblicazione: (2023)
di: Rothenberger, Jay C., et al.
Pubblicazione: (2023)
Co-Training with Active Contrastive Learning and Meta-Pseudo-Labeling on 2D Projections for Deep Semi-Supervised Learning
di: Aparco-Cardenas, David, et al.
Pubblicazione: (2025)
di: Aparco-Cardenas, David, et al.
Pubblicazione: (2025)
VLSlice: Interactive Vision-and-Language Slice Discovery
di: Slyman, Eric, et al.
Pubblicazione: (2023)
di: Slyman, Eric, et al.
Pubblicazione: (2023)
Towards Onboard Continuous Change Detection for Floods
di: Kyselica, Daniel, et al.
Pubblicazione: (2026)
di: Kyselica, Daniel, et al.
Pubblicazione: (2026)
Non-Robust Features are Not Always Useful in One-Class Classification
di: Lau, Matthew, et al.
Pubblicazione: (2024)
di: Lau, Matthew, et al.
Pubblicazione: (2024)
Parameterizing Dataset Distillation via Gaussian Splatting
di: Jiang, Chenyang, et al.
Pubblicazione: (2025)
di: Jiang, Chenyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Unveiling Text in Challenging Stone Inscriptions: A Character-Context-Aware Patching Strategy for Binarization
di: Jena, Pratyush, et al.
Pubblicazione: (2026) -
LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
di: Jia, Yongju, et al.
Pubblicazione: (2025) -
A Spitting Image: Modular Superpixel Tokenization in Vision Transformers
di: Aasan, Marius, et al.
Pubblicazione: (2024) -
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
di: Yasuno, Takato
Pubblicazione: (2026) -
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
di: Chaybouti, Sofian, et al.
Pubblicazione: (2025)