Revisiting [CLS] and Patch Token Interaction in Vision Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Marouani, Alexis, Siméoni, Oriane, Jégou, Hervé, Bojanowski, Piotr, Vo, Huy V. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unveiling Text in Challenging Stone Inscriptions: A Character-Context-Aware Patching Strategy for Binarization
von: Jena, Pratyush, et al.
Veröffentlicht: (2026)
von: Jena, Pratyush, et al.
Veröffentlicht: (2026)
LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
von: Jia, Yongju, et al.
Veröffentlicht: (2025)
von: Jia, Yongju, et al.
Veröffentlicht: (2025)
A Spitting Image: Modular Superpixel Tokenization in Vision Transformers
von: Aasan, Marius, et al.
Veröffentlicht: (2024)
von: Aasan, Marius, et al.
Veröffentlicht: (2024)
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
von: Yasuno, Takato
Veröffentlicht: (2026)
von: Yasuno, Takato
Veröffentlicht: (2026)
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)
Learning Unified Representation of 3D Gaussian Splatting
von: Xin, Yuelin, et al.
Veröffentlicht: (2025)
von: Xin, Yuelin, et al.
Veröffentlicht: (2025)
VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
HySparK: Hybrid Sparse Masking for Large Scale Medical Image Pre-Training
von: Tang, Fenghe, et al.
Veröffentlicht: (2024)
von: Tang, Fenghe, et al.
Veröffentlicht: (2024)
Mobile-Ready Automated Triage of Diabetic Retinopathy Using Digital Fundus Images
von: Joshi, Aadi, et al.
Veröffentlicht: (2026)
von: Joshi, Aadi, et al.
Veröffentlicht: (2026)
Differentiable Hierarchical Visual Tokenization
von: Aasan, Marius, et al.
Veröffentlicht: (2025)
von: Aasan, Marius, et al.
Veröffentlicht: (2025)
Frequency-Decomposed INR for NIR-Assisted Low-Light RGB Image Denoising
von: Shi, Ligen, et al.
Veröffentlicht: (2026)
von: Shi, Ligen, et al.
Veröffentlicht: (2026)
Neural Fields for 3D Tracking of Anatomy and Surgical Instruments in Monocular Laparoscopic Video Clips
von: Gerats, Beerend G. A., et al.
Veröffentlicht: (2024)
von: Gerats, Beerend G. A., et al.
Veröffentlicht: (2024)
A Hierarchical Self-Consistent Regularization Approach to Satellite Image Time Series Classification
von: Weikmann, Giulio, et al.
Veröffentlicht: (2025)
von: Weikmann, Giulio, et al.
Veröffentlicht: (2025)
Learning to Expand Images for Efficient Visual Autoregressive Modeling
von: Yang, Ruiqing, et al.
Veröffentlicht: (2025)
von: Yang, Ruiqing, et al.
Veröffentlicht: (2025)
Scalable Face Security Vision Foundation Model for Deepfake, Diffusion, and Spoofing Detection
von: Wang, Gaojian, et al.
Veröffentlicht: (2025)
von: Wang, Gaojian, et al.
Veröffentlicht: (2025)
Cora: Correspondence-aware image editing using few step diffusion
von: Alimohammadi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Alimohammadi, Amirhossein, et al.
Veröffentlicht: (2025)
Pointing-Based Object Recognition
von: Hajdúch, Lukáš, et al.
Veröffentlicht: (2026)
von: Hajdúch, Lukáš, et al.
Veröffentlicht: (2026)
Prototype-Guided Concept Erasure in Diffusion Models
von: Cai, Yuze, et al.
Veröffentlicht: (2026)
von: Cai, Yuze, et al.
Veröffentlicht: (2026)
Supervised Contrastive Learning for Few-Shot AI-Generated Image Detection and Attribution
von: Urueña, Jaime Álvarez, et al.
Veröffentlicht: (2025)
von: Urueña, Jaime Álvarez, et al.
Veröffentlicht: (2025)
Label Delay in Online Continual Learning
von: Csaba, Botos, et al.
Veröffentlicht: (2023)
von: Csaba, Botos, et al.
Veröffentlicht: (2023)
Neural Implicit Morphing of Face Images
von: Schardong, Guilherme, et al.
Veröffentlicht: (2023)
von: Schardong, Guilherme, et al.
Veröffentlicht: (2023)
RealHD: A High-Quality Dataset for Robust Detection of State-of-the-Art AI-Generated Images
von: Yu, Hanzhe, et al.
Veröffentlicht: (2026)
von: Yu, Hanzhe, et al.
Veröffentlicht: (2026)
A Novel Global Context-aware Deep Neural Network for Enhanced Brain Tumor Segmentation using Magnetic Resonance Images
von: Mukherjee, Sourjya, et al.
Veröffentlicht: (2026)
von: Mukherjee, Sourjya, et al.
Veröffentlicht: (2026)
When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation
von: Frochte, Jörg
Veröffentlicht: (2026)
von: Frochte, Jörg
Veröffentlicht: (2026)
FAME: Feature Activation Map Explanation on Image Classification and Face Recognition
von: Zhang, Xinyi, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyi, et al.
Veröffentlicht: (2026)
UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis Augmentation
von: Zhao, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhao, Zhihao, et al.
Veröffentlicht: (2025)
GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations
von: Liu, Zeping, et al.
Veröffentlicht: (2025)
von: Liu, Zeping, et al.
Veröffentlicht: (2025)
MienCap: Realtime Performance-Based Facial Animation with Live Mood Dynamics
von: Pan, Ye, et al.
Veröffentlicht: (2025)
von: Pan, Ye, et al.
Veröffentlicht: (2025)
Generating real-time detailed ground visualisations from sparse aerial point clouds
von: Murray, Aidan, et al.
Veröffentlicht: (2025)
von: Murray, Aidan, et al.
Veröffentlicht: (2025)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
Symmetry Awareness Encoded Deep Learning Framework for Brain Imaging Analysis
von: Ma, Yang, et al.
Veröffentlicht: (2024)
von: Ma, Yang, et al.
Veröffentlicht: (2024)
STimage-1K4M: A histopathology image-gene expression dataset for spatial transcriptomics
von: Chen, Jiawen, et al.
Veröffentlicht: (2024)
von: Chen, Jiawen, et al.
Veröffentlicht: (2024)
ViFiCon: Vision and Wireless Association Via Self-Supervised Contrastive Learning
von: Meegan, Nicholas, et al.
Veröffentlicht: (2022)
von: Meegan, Nicholas, et al.
Veröffentlicht: (2022)
Implicitly Learned Neural Phase Functions for Basis-Free Point Spread Function Engineering
von: Valouev, Aleksey
Veröffentlicht: (2024)
von: Valouev, Aleksey
Veröffentlicht: (2024)
Meta Co-Training: Two Views are Better than One
von: Rothenberger, Jay C., et al.
Veröffentlicht: (2023)
von: Rothenberger, Jay C., et al.
Veröffentlicht: (2023)
Co-Training with Active Contrastive Learning and Meta-Pseudo-Labeling on 2D Projections for Deep Semi-Supervised Learning
von: Aparco-Cardenas, David, et al.
Veröffentlicht: (2025)
von: Aparco-Cardenas, David, et al.
Veröffentlicht: (2025)
VLSlice: Interactive Vision-and-Language Slice Discovery
von: Slyman, Eric, et al.
Veröffentlicht: (2023)
von: Slyman, Eric, et al.
Veröffentlicht: (2023)
Towards Onboard Continuous Change Detection for Floods
von: Kyselica, Daniel, et al.
Veröffentlicht: (2026)
von: Kyselica, Daniel, et al.
Veröffentlicht: (2026)
Non-Robust Features are Not Always Useful in One-Class Classification
von: Lau, Matthew, et al.
Veröffentlicht: (2024)
von: Lau, Matthew, et al.
Veröffentlicht: (2024)
Parameterizing Dataset Distillation via Gaussian Splatting
von: Jiang, Chenyang, et al.
Veröffentlicht: (2025)
von: Jiang, Chenyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unveiling Text in Challenging Stone Inscriptions: A Character-Context-Aware Patching Strategy for Binarization
von: Jena, Pratyush, et al.
Veröffentlicht: (2026) -
LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
von: Jia, Yongju, et al.
Veröffentlicht: (2025) -
A Spitting Image: Modular Superpixel Tokenization in Vision Transformers
von: Aasan, Marius, et al.
Veröffentlicht: (2024) -
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
von: Yasuno, Takato
Veröffentlicht: (2026) -
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)