Two-stage Vision Transformers and Hard Masking offer Robust Object Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Aniraj, Ananthu, Dantas, Cassio F., Ienco, Dino, Marcos, Diego |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PDiscoFormer: Relaxing Part Discovery Constraints with Vision Transformers
by: Aniraj, Ananthu, et al.
Published: (2024)
by: Aniraj, Ananthu, et al.
Published: (2024)
Metonymy in vision models undermines attention-based interpretability
by: Aniraj, Ananthu, et al.
Published: (2026)
by: Aniraj, Ananthu, et al.
Published: (2026)
DisCoM-KD: Cross-Modal Knowledge Distillation via Disentanglement Representation and Adversarial Learning
by: Ienco, Dino, et al.
Published: (2024)
by: Ienco, Dino, et al.
Published: (2024)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025)
by: Mena, Francisco, et al.
Published: (2025)
Revisiting Cross-Modal Knowledge Distillation: A Disentanglement Approach for RGBD Semantic Segmentation
by: Ferrod, Roger, et al.
Published: (2025)
by: Ferrod, Roger, et al.
Published: (2025)
Towards Robust Vision Transformer via Masked Adaptive Ensemble
by: Lin, Fudong, et al.
Published: (2024)
by: Lin, Fudong, et al.
Published: (2024)
TimeSenCLIP: A Time Series Vision-Language Model for Remote Sensing
by: Jain, Pallavi, et al.
Published: (2025)
by: Jain, Pallavi, et al.
Published: (2025)
RTAT: A Robust Two-stage Association Tracker for Multi-Object Tracking
by: Guo, Song, et al.
Published: (2024)
by: Guo, Song, et al.
Published: (2024)
Masked Modeling for Self-supervised Representation Learning on Vision and Beyond
by: Li, Siyuan, et al.
Published: (2023)
by: Li, Siyuan, et al.
Published: (2023)
Tackling the Abstraction and Reasoning Corpus with Vision Transformers: the Importance of 2D Representation, Positions, and Objects
by: Li, Wenhao, et al.
Published: (2024)
by: Li, Wenhao, et al.
Published: (2024)
Proximal Vision Transformer: Enhancing Feature Representation through Two-Stage Manifold Geometry
by: Yun, Haoyu, et al.
Published: (2025)
by: Yun, Haoyu, et al.
Published: (2025)
Salient Mask-Guided Vision Transformer for Fine-Grained Classification
by: Demidov, Dmitry, et al.
Published: (2023)
by: Demidov, Dmitry, et al.
Published: (2023)
SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting
by: Jain, Pallavi, et al.
Published: (2024)
by: Jain, Pallavi, et al.
Published: (2024)
Human-Like Coarse Object Representations in Vision Models
by: Gizdov, Andrey, et al.
Published: (2026)
by: Gizdov, Andrey, et al.
Published: (2026)
Representation Separation for Semantic Segmentation with Vision Transformers
by: Hong, Yuanduo, et al.
Published: (2022)
by: Hong, Yuanduo, et al.
Published: (2022)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained Clients
by: Wu, Meihan, et al.
Published: (2024)
by: Wu, Meihan, et al.
Published: (2024)
MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
Cross-Task Multi-Branch Vision Transformer for Facial Expression and Mask Wearing Classification
by: Zhu, Armando, et al.
Published: (2024)
by: Zhu, Armando, et al.
Published: (2024)
Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
Hybrid Spiking Vision Transformer for Object Detection with Event Cameras
by: Xu, Qi, et al.
Published: (2025)
by: Xu, Qi, et al.
Published: (2025)
Robust Multiple Description Neural Video Codec with Masked Transformer for Dynamic and Noisy Networks
by: Hu, Xinyue, et al.
Published: (2024)
by: Hu, Xinyue, et al.
Published: (2024)
GS2Pose: Two-stage 6D Object Pose Estimation Guided by Gaussian Splatting
by: Mei, Jilan, et al.
Published: (2024)
by: Mei, Jilan, et al.
Published: (2024)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
by: Baraldi, Lorenzo, et al.
Published: (2023)
by: Baraldi, Lorenzo, et al.
Published: (2023)
Semi Supervised Heterogeneous Domain Adaptation via Disentanglement and Pseudo-Labelling
by: Dantas, Cassio F., et al.
Published: (2024)
by: Dantas, Cassio F., et al.
Published: (2024)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
by: Lepori, Michael A., et al.
Published: (2024)
by: Lepori, Michael A., et al.
Published: (2024)
Mask-RadarNet: Enhancing Transformer With Spatial-Temporal Semantic Context for Radar Object Detection in Autonomous Driving
by: Wu, Yuzhi, et al.
Published: (2024)
by: Wu, Yuzhi, et al.
Published: (2024)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
by: Song, Chull Hwan, et al.
Published: (2024)
by: Song, Chull Hwan, et al.
Published: (2024)
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
by: Vani, Ankit, et al.
Published: (2024)
by: Vani, Ankit, et al.
Published: (2024)
The Progression of Transformers from Language to Vision to MOT: A Literature Review on Multi-Object Tracking with Transformers
by: Kamboj, Abhi
Published: (2024)
by: Kamboj, Abhi
Published: (2024)
COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
by: Li, Wenxi, et al.
Published: (2025)
by: Li, Wenxi, et al.
Published: (2025)
Masked Latent Transformer with the Random Masking Ratio to Advance the Diagnosis of Dental Fluorosis
by: Wu, Yun, et al.
Published: (2024)
by: Wu, Yun, et al.
Published: (2024)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
by: Gandhi, Sanket, et al.
Published: (2024)
by: Gandhi, Sanket, et al.
Published: (2024)
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
by: Shi, Dai
Published: (2023)
by: Shi, Dai
Published: (2023)
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
by: Shams, Montasir, et al.
Published: (2025)
by: Shams, Montasir, et al.
Published: (2025)
Symbolic Rule Extraction from Attention-Guided Sparse Representations in Vision Transformers
by: Padalkar, Parth, et al.
Published: (2025)
by: Padalkar, Parth, et al.
Published: (2025)
IFViT: Interpretable Fixed-Length Representation for Fingerprint Matching via Vision Transformer
by: Qiu, Yuhang, et al.
Published: (2024)
by: Qiu, Yuhang, et al.
Published: (2024)
I Walk the Line: Examining the Role of Gestalt Continuity in Object Binding for Vision Transformers
by: Tartaglini, Alexa R., et al.
Published: (2026)
by: Tartaglini, Alexa R., et al.
Published: (2026)
Similar Items
-
PDiscoFormer: Relaxing Part Discovery Constraints with Vision Transformers
by: Aniraj, Ananthu, et al.
Published: (2024) -
Metonymy in vision models undermines attention-based interpretability
by: Aniraj, Ananthu, et al.
Published: (2026) -
DisCoM-KD: Cross-Modal Knowledge Distillation via Disentanglement Representation and Adversarial Learning
by: Ienco, Dino, et al.
Published: (2024) -
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025) -
Revisiting Cross-Modal Knowledge Distillation: A Disentanglement Approach for RGBD Semantic Segmentation
by: Ferrod, Roger, et al.
Published: (2025)