PDiscoFormer: Relaxing Part Discovery Constraints with Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Aniraj, Ananthu, Dantas, Cassio F., Ienco, Dino, Marcos, Diego |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Two-stage Vision Transformers and Hard Masking offer Robust Object Representations
by: Aniraj, Ananthu, et al.
Published: (2025)
by: Aniraj, Ananthu, et al.
Published: (2025)
DisCoM-KD: Cross-Modal Knowledge Distillation via Disentanglement Representation and Adversarial Learning
by: Ienco, Dino, et al.
Published: (2024)
by: Ienco, Dino, et al.
Published: (2024)
Metonymy in vision models undermines attention-based interpretability
by: Aniraj, Ananthu, et al.
Published: (2026)
by: Aniraj, Ananthu, et al.
Published: (2026)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025)
by: Mena, Francisco, et al.
Published: (2025)
Semi Supervised Heterogeneous Domain Adaptation via Disentanglement and Pseudo-Labelling
by: Dantas, Cassio F., et al.
Published: (2024)
by: Dantas, Cassio F., et al.
Published: (2024)
MetaFormer Baselines for Vision
by: Yu, Weihao, et al.
Published: (2022)
by: Yu, Weihao, et al.
Published: (2022)
Geographical Context Matters: Bridging Fine and Coarse Spatial Information to Enhance Continental Land Cover Mapping
by: Ghassemi, Babak, et al.
Published: (2025)
by: Ghassemi, Babak, et al.
Published: (2025)
iFormer: Integrating ConvNet and Transformer for Mobile Application
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
Towards a multimodal framework for remote sensing image change retrieval and captioning
by: Ferrod, Roger, et al.
Published: (2024)
by: Ferrod, Roger, et al.
Published: (2024)
Revisiting Cross-Modal Knowledge Distillation: A Disentanglement Approach for RGBD Semantic Segmentation
by: Ferrod, Roger, et al.
Published: (2025)
by: Ferrod, Roger, et al.
Published: (2025)
HGTS-Former: Hierarchical HyperGraph Transformer for Multivariate Time Series Analysis
by: Si, Hao, et al.
Published: (2025)
by: Si, Hao, et al.
Published: (2025)
ScribFormer: Transformer Makes CNN Work Better for Scribble-based Medical Image Segmentation
by: Li, Zihan, et al.
Published: (2024)
by: Li, Zihan, et al.
Published: (2024)
SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and $\mathcal{O}(T)$ Complexity
by: Zou, Shihao, et al.
Published: (2025)
by: Zou, Shihao, et al.
Published: (2025)
RigidFormer: Learning Rigid Dynamics using Transformers
by: Dou, Zhiyang, et al.
Published: (2026)
by: Dou, Zhiyang, et al.
Published: (2026)
Soft-TransFormers for Continual Learning
by: Kang, Haeyong, et al.
Published: (2024)
by: Kang, Haeyong, et al.
Published: (2024)
Constraint-Aware Neurosymbolic Uncertainty Quantification with Bayesian Deep Learning for Scientific Discovery
by: Alam, Shahnawaz, et al.
Published: (2026)
by: Alam, Shahnawaz, et al.
Published: (2026)
Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models
by: Wang, Hongjun, et al.
Published: (2026)
by: Wang, Hongjun, et al.
Published: (2026)
MDP: Multidimensional Vision Model Pruning with Latency Constraint
by: Sun, Xinglong, et al.
Published: (2025)
by: Sun, Xinglong, et al.
Published: (2025)
Learning Part Knowledge to Facilitate Category Understanding for Fine-Grained Generalized Category Discovery
by: Wang, Enguang, et al.
Published: (2025)
by: Wang, Enguang, et al.
Published: (2025)
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)
by: Hu, Lijie, et al.
Published: (2023)
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025)
by: Jacobs, Mozes, et al.
Published: (2025)
JetFormer: An Autoregressive Generative Model of Raw Images and Text
by: Tschannen, Michael, et al.
Published: (2024)
by: Tschannen, Michael, et al.
Published: (2024)
Continual Adaptation of Vision Transformers for Federated Learning
by: Halbe, Shaunak, et al.
Published: (2023)
by: Halbe, Shaunak, et al.
Published: (2023)
Mechanisms of Non-Monotonic Scaling in Vision Transformers
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
Discovering Influential Neuron Path in Vision Transformers
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
DiffiT: Diffusion Vision Transformers for Image Generation
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
ADAPT to Robustify Prompt Tuning Vision Transformers
by: Eskandar, Masih, et al.
Published: (2024)
by: Eskandar, Masih, et al.
Published: (2024)
Accelerating Vision Transformers with Adaptive Patch Sizes
by: Choudhury, Rohan, et al.
Published: (2025)
by: Choudhury, Rohan, et al.
Published: (2025)
Class-Discriminative Attention Maps for Vision Transformers
by: Brocki, Lennart, et al.
Published: (2023)
by: Brocki, Lennart, et al.
Published: (2023)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)
by: Fiastre, Gabriel, et al.
Published: (2025)
CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
by: Nguyen, Kiet A., et al.
Published: (2024)
by: Nguyen, Kiet A., et al.
Published: (2024)
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
Oscillation-Reduced MXFP4 Training for Vision Transformers
by: Chen, Yuxiang, et al.
Published: (2025)
by: Chen, Yuxiang, et al.
Published: (2025)
Enhancing Vision Transformer Explainability Using Artificial Astrocytes
by: Echevarrieta-Catalan, Nicolas, et al.
Published: (2025)
by: Echevarrieta-Catalan, Nicolas, et al.
Published: (2025)
FasterViT: Fast Vision Transformers with Hierarchical Attention
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
Understanding Video Transformers via Universal Concept Discovery
by: Kowal, Matthew, et al.
Published: (2024)
by: Kowal, Matthew, et al.
Published: (2024)
MindFormer: Semantic Alignment of Multi-Subject fMRI for Brain Decoding
by: Han, Inhwa, et al.
Published: (2024)
by: Han, Inhwa, et al.
Published: (2024)
LucidPPN: Unambiguous Prototypical Parts Network for User-centric Interpretable Computer Vision
by: Pach, Mateusz, et al.
Published: (2024)
by: Pach, Mateusz, et al.
Published: (2024)
VariViT: A Vision Transformer for Variable Image Sizes
by: Varma, Aswathi, et al.
Published: (2026)
by: Varma, Aswathi, et al.
Published: (2026)
A Survey of the Self Supervised Learning Mechanisms for Vision Transformers
by: Khan, Asifullah, et al.
Published: (2024)
by: Khan, Asifullah, et al.
Published: (2024)
Similar Items
-
Two-stage Vision Transformers and Hard Masking offer Robust Object Representations
by: Aniraj, Ananthu, et al.
Published: (2025) -
DisCoM-KD: Cross-Modal Knowledge Distillation via Disentanglement Representation and Adversarial Learning
by: Ienco, Dino, et al.
Published: (2024) -
Metonymy in vision models undermines attention-based interpretability
by: Aniraj, Ananthu, et al.
Published: (2026) -
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025) -
Semi Supervised Heterogeneous Domain Adaptation via Disentanglement and Pseudo-Labelling
by: Dantas, Cassio F., et al.
Published: (2024)