A Closer Look at Multimodal Representation Collapse
Fuente:
arXiv
Saved in:
| Main Authors: | Chaudhuri, Abhra, Dutta, Anjan, Bui, Tu, Georgescu, Serban |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Conditional Invariances through Non-Commutativity
by: Chaudhuri, Abhra, et al.
Published: (2024)
by: Chaudhuri, Abhra, et al.
Published: (2024)
DeNetDM: Debiasing by Network Depth Modulation
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2024)
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2024)
Scalable and Loosely-Coupled Multimodal Deep Learning for Breast Cancer Subtyping
by: Amer, Mohammed, et al.
Published: (2025)
by: Amer, Mohammed, et al.
Published: (2025)
CLIPDraw++: Text-to-Sketch Synthesis with Simple Primitives
by: Mathur, Nityanand, et al.
Published: (2023)
by: Mathur, Nityanand, et al.
Published: (2023)
A Closer Look at Data Augmentation Strategies for Finetuning-Based Low/Few-Shot Object Detection
by: Li, Vladislav, et al.
Published: (2024)
by: Li, Vladislav, et al.
Published: (2024)
Simplifying Multi-Task Architectures Through Task-Specific Normalization
by: Suteu, Mihai, et al.
Published: (2025)
by: Suteu, Mihai, et al.
Published: (2025)
MaxSup: Overcoming Representation Collapse in Label Smoothing
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Subspace-Boosted Model Merging
by: Skorobogat, Ronald, et al.
Published: (2025)
by: Skorobogat, Ronald, et al.
Published: (2025)
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
by: Lin, Han, et al.
Published: (2026)
by: Lin, Han, et al.
Published: (2026)
Signal Collapse in One-Shot Pruning: When Sparse Models Fail to Distinguish Neural Representations
by: Saikumar, Dhananjay, et al.
Published: (2025)
by: Saikumar, Dhananjay, et al.
Published: (2025)
Gramian Multimodal Representation Learning and Alignment
by: Cicchetti, Giordano, et al.
Published: (2024)
by: Cicchetti, Giordano, et al.
Published: (2024)
Actor-agnostic Multi-label Action Recognition with Multi-modal Query
by: Mondal, Anindya, et al.
Published: (2023)
by: Mondal, Anindya, et al.
Published: (2023)
Align Your Query: Representation Alignment for Multimodality Medical Object Detection
by: Seo, Ara, et al.
Published: (2025)
by: Seo, Ara, et al.
Published: (2025)
Invariant Representation Guided Multimodal Sentiment Decoding with Sequential Variation Regularization
by: Xu, Guoyang, et al.
Published: (2024)
by: Xu, Guoyang, et al.
Published: (2024)
Learning Relative Representations for Fine-Grained Multimodal Alignment with Limited Data
by: Kim, Shiwon, et al.
Published: (2026)
by: Kim, Shiwon, et al.
Published: (2026)
A Fresh Look at Sanity Checks for Saliency Maps
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
by: Yang, Yaoxin, et al.
Published: (2025)
by: Yang, Yaoxin, et al.
Published: (2025)
NECO: NEural Collapse Based Out-of-distribution detection
by: Ammar, Mouïn Ben, et al.
Published: (2023)
by: Ammar, Mouïn Ben, et al.
Published: (2023)
Memory-efficient Continual Learning with Neural Collapse Contrastive
by: Dang, Trung-Anh, et al.
Published: (2024)
by: Dang, Trung-Anh, et al.
Published: (2024)
Less is More: A Closer Look at Semantic-based Few-Shot Learning
by: Zhou, Chunpeng, et al.
Published: (2024)
by: Zhou, Chunpeng, et al.
Published: (2024)
Adapting Vision-Language Models for Evaluating World Models
by: Hendriksen, Mariya, et al.
Published: (2025)
by: Hendriksen, Mariya, et al.
Published: (2025)
Controllable Lung Nodule Synthesis via Histogram-Regularized Latent Diffusion Models
by: Kannan, Arunkumar, et al.
Published: (2026)
by: Kannan, Arunkumar, et al.
Published: (2026)
On Partial Prototype Collapse in the DINO Family of Self-Supervised Methods
by: Govindarajan, Hariprasath, et al.
Published: (2024)
by: Govindarajan, Hariprasath, et al.
Published: (2024)
Learning Noise-Robust Joint Representation for Multimodal Emotion Recognition under Incomplete Data Scenarios
by: Fan, Qi, et al.
Published: (2023)
by: Fan, Qi, et al.
Published: (2023)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
by: Kim, Sanghwan, et al.
Published: (2024)
by: Kim, Sanghwan, et al.
Published: (2024)
Weight Copy and Low-Rank Adaptation for Few-Shot Distillation of Vision Transformers
by: Grigore, Diana-Nicoleta, et al.
Published: (2024)
by: Grigore, Diana-Nicoleta, et al.
Published: (2024)
TopoFR: A Closer Look at Topology Alignment on Face Recognition
by: Dan, Jun, et al.
Published: (2024)
by: Dan, Jun, et al.
Published: (2024)
Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making
by: Greer, Ross, et al.
Published: (2026)
by: Greer, Ross, et al.
Published: (2026)
Temporal Embeddings: Scalable Self-Supervised Temporal Representation Learning from Spatiotemporal Data for Multimodal Computer Vision
by: Cao, Yi, et al.
Published: (2023)
by: Cao, Yi, et al.
Published: (2023)
Spherical VAE with Cluster-Aware Feasible Regions: Guaranteed Prevention of Posterior Collapse
by: Zhang, Zegu, et al.
Published: (2026)
by: Zhang, Zegu, et al.
Published: (2026)
A Gray-box Attack against Latent Diffusion Model-based Image Editing by Posterior Collapse
by: Guo, Zhongliang, et al.
Published: (2024)
by: Guo, Zhongliang, et al.
Published: (2024)
A Markovian View of Iterative-Feedback Loops in Image Generative Models: Neural Resonance and Model Collapse
by: Vats, Vibhas Kumar, et al.
Published: (2026)
by: Vats, Vibhas Kumar, et al.
Published: (2026)
This Looks Better than That: Better Interpretable Models with ProtoPNeXt
by: Willard, Frank, et al.
Published: (2024)
by: Willard, Frank, et al.
Published: (2024)
Healthy Harvests: A Comparative Look at Guava Disease Classification Using InceptionV3
by: Ghosh, Samanta, et al.
Published: (2026)
by: Ghosh, Samanta, et al.
Published: (2026)
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
by: Zhou, Yue, et al.
Published: (2026)
by: Zhou, Yue, et al.
Published: (2026)
Mitigating Catastrophic Forgetting and Mode Collapse in Text-to-Image Diffusion via Latent Replay
by: Otani, Aoi
Published: (2025)
by: Otani, Aoi
Published: (2025)
Look Through Masks: Towards Masked Face Recognition with De-Occlusion Distillation
by: Li, Chenyu, et al.
Published: (2024)
by: Li, Chenyu, et al.
Published: (2024)
Learning to Infer Generative Template Programs for Visual Concepts
by: Jones, R. Kenny, et al.
Published: (2024)
by: Jones, R. Kenny, et al.
Published: (2024)
Similar Items
-
Learning Conditional Invariances through Non-Commutativity
by: Chaudhuri, Abhra, et al.
Published: (2024) -
DeNetDM: Debiasing by Network Depth Modulation
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2024) -
Scalable and Loosely-Coupled Multimodal Deep Learning for Breast Cancer Subtyping
by: Amer, Mohammed, et al.
Published: (2025) -
CLIPDraw++: Text-to-Sketch Synthesis with Simple Primitives
by: Mathur, Nityanand, et al.
Published: (2023) -
A Closer Look at Data Augmentation Strategies for Finetuning-Based Low/Few-Shot Object Detection
by: Li, Vladislav, et al.
Published: (2024)