Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Koepke, A. Sophia, Zverev, Daniil, Ginosar, Shiry, Efros, Alexei A. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
by: Zverev, Daniil, et al.
Published: (2025)
by: Zverev, Daniil, et al.
Published: (2025)
Diffusion Models as Data Mining Tools
by: Siglidis, Ioannis, et al.
Published: (2024)
by: Siglidis, Ioannis, et al.
Published: (2024)
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
by: Harrington, Anne, et al.
Published: (2025)
by: Harrington, Anne, et al.
Published: (2025)
Frozen Forecasting: A Unified Evaluation
by: Walker, Jacob C, et al.
Published: (2025)
by: Walker, Jacob C, et al.
Published: (2025)
KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models
by: Yiu, Eunice, et al.
Published: (2024)
by: Yiu, Eunice, et al.
Published: (2024)
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)
by: Gandelsman, Yossi, et al.
Published: (2023)
VGGSounder: Audio-Visual Evaluations for Foundation Models
by: Zverev, Daniil, et al.
Published: (2025)
by: Zverev, Daniil, et al.
Published: (2025)
Gaussian Masked Autoencoders
by: Rajasegaran, Jathushan, et al.
Published: (2025)
by: Rajasegaran, Jathushan, et al.
Published: (2025)
SpectralGCD: Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
by: Caselli, Lorenzo, et al.
Published: (2026)
by: Caselli, Lorenzo, et al.
Published: (2026)
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
by: Camuffo, Elena, et al.
Published: (2025)
by: Camuffo, Elena, et al.
Published: (2025)
Scaling 4D Representations
by: Carreira, João, et al.
Published: (2024)
by: Carreira, João, et al.
Published: (2024)
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
by: Mistretta, Marco, et al.
Published: (2025)
by: Mistretta, Marco, et al.
Published: (2025)
CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly Detection
by: Wang, Xiaolei, et al.
Published: (2024)
by: Wang, Xiaolei, et al.
Published: (2024)
Explore the Limits of Omni-modal Pretraining at Scale
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Multi-modal Masked Siamese Network Improves Chest X-Ray Representation Learning
by: Shurrab, Saeed, et al.
Published: (2024)
by: Shurrab, Saeed, et al.
Published: (2024)
Escaping Plato's Cave: JAM for Aligning Independently Trained Vision and Language Models
by: Yoon, Lauren Hyoseo, et al.
Published: (2025)
by: Yoon, Lauren Hyoseo, et al.
Published: (2025)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025)
by: Mena, Francisco, et al.
Published: (2025)
Cross-modal Affinity-aligned Multimodal Learning Analytics for Predicting Student Collaboration Satisfaction in Game-Based Learning
by: Tsai, Wen-Hsin, et al.
Published: (2026)
by: Tsai, Wen-Hsin, et al.
Published: (2026)
Text-Guided Multi-Scale Frequency Representation Adaptation
by: Yan, Weicai, et al.
Published: (2026)
by: Yan, Weicai, et al.
Published: (2026)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
AmCLR: Unified Augmented Learning for Cross-Modal Representations
by: Jagannath, Ajay, et al.
Published: (2024)
by: Jagannath, Ajay, et al.
Published: (2024)
Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data
by: Yoon, Heegeon, et al.
Published: (2026)
by: Yoon, Heegeon, et al.
Published: (2026)
GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs
by: Hua, Pu, et al.
Published: (2024)
by: Hua, Pu, et al.
Published: (2024)
Scratching Visual Transformer's Back with Uniform Attention
by: Hyeon-Woo, Nam, et al.
Published: (2022)
by: Hyeon-Woo, Nam, et al.
Published: (2022)
Cross-modal linkage risk in clinical vision-language models
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
A General Framework for Robust G-Invariance in G-Equivariant Networks
by: Sanborn, Sophia, et al.
Published: (2023)
by: Sanborn, Sophia, et al.
Published: (2023)
Scaling Up Your Kernels: Large Kernel Design in ConvNets towards Universal Representations
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
DisCoM-KD: Cross-Modal Knowledge Distillation via Disentanglement Representation and Adversarial Learning
by: Ienco, Dino, et al.
Published: (2024)
by: Ienco, Dino, et al.
Published: (2024)
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
by: Cherepanov, Egor, et al.
Published: (2026)
by: Cherepanov, Egor, et al.
Published: (2026)
Object-Centric Learning with Slot Mixture Module
by: Kirilenko, Daniil, et al.
Published: (2023)
by: Kirilenko, Daniil, et al.
Published: (2023)
Examining Common Paradigms in Multi-Task Learning
by: Elich, Cathrin, et al.
Published: (2023)
by: Elich, Cathrin, et al.
Published: (2023)
Visual Hallucinations of Multi-modal Large Language Models
by: Huang, Wen, et al.
Published: (2024)
by: Huang, Wen, et al.
Published: (2024)
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
by: Zhu, Mengdan, et al.
Published: (2025)
by: Zhu, Mengdan, et al.
Published: (2025)
Generative Multi-modal Models are Good Class-Incremental Learners
by: Cao, Xusheng, et al.
Published: (2024)
by: Cao, Xusheng, et al.
Published: (2024)
Multi-modal Vision Pre-training for Medical Image Analysis
by: Rui, Shaohao, et al.
Published: (2024)
by: Rui, Shaohao, et al.
Published: (2024)
X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
by: Wang, Yunzhe, et al.
Published: (2025)
by: Wang, Yunzhe, et al.
Published: (2025)
Back Home: A Computer Vision Solution to Seashell Identification for Ecological Restoration
by: Valverde, Alexander, et al.
Published: (2025)
by: Valverde, Alexander, et al.
Published: (2025)
Disentangled 3D Scene Generation with Layout Learning
by: Epstein, Dave, et al.
Published: (2024)
by: Epstein, Dave, et al.
Published: (2024)
MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
Similar Items
-
On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
by: Zverev, Daniil, et al.
Published: (2025) -
Diffusion Models as Data Mining Tools
by: Siglidis, Ioannis, et al.
Published: (2024) -
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
by: Harrington, Anne, et al.
Published: (2025) -
Frozen Forecasting: A Unified Evaluation
by: Walker, Jacob C, et al.
Published: (2025) -
KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models
by: Yiu, Eunice, et al.
Published: (2024)