Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Sharut, Sundaram, Shobhita, Wang, Chenyu, Jegelka, Stefanie, Isola, Phillip |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
In-Context Symmetries: Self-Supervised Learning through Contextual World Models
by: Gupta, Sharut, et al.
Published: (2024)
by: Gupta, Sharut, et al.
Published: (2024)
Personalized Representation from Personalized Generation
by: Sundaram, Shobhita, et al.
Published: (2024)
by: Sundaram, Shobhita, et al.
Published: (2024)
DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data
by: Fu, Stephanie, et al.
Published: (2023)
by: Fu, Stephanie, et al.
Published: (2023)
Learning Diffusion Models with Flexible Representation Guidance
by: Wang, Chenyu, et al.
Published: (2025)
by: Wang, Chenyu, et al.
Published: (2025)
When Does Perceptual Alignment Benefit Vision Representations?
by: Sundaram, Shobhita, et al.
Published: (2024)
by: Sundaram, Shobhita, et al.
Published: (2024)
Understanding the Role of Equivariance in Self-supervised Learning
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Canonicalizing Multimodal Contrastive Representation Learning
by: Gupta, Sharut, et al.
Published: (2026)
by: Gupta, Sharut, et al.
Published: (2026)
Words That Make Language Models Perceive
by: Wang, Sophie L., et al.
Published: (2025)
by: Wang, Sophie L., et al.
Published: (2025)
Better Together: Evaluating the Complementarity of Earth Embedding Models
by: van der Plas, Thijs L, et al.
Published: (2026)
by: van der Plas, Thijs L, et al.
Published: (2026)
Multimodal Representation Learning by Alternating Unimodal Adaptation
by: Zhang, Xiaohui, et al.
Published: (2023)
by: Zhang, Xiaohui, et al.
Published: (2023)
Backdoor Attack on Unpaired Medical Image-Text Foundation Models: A Pilot Study on MedCLIP
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
by: Bahng, Hyojin, et al.
Published: (2025)
by: Bahng, Hyojin, et al.
Published: (2025)
Learning Shared Representations from Unpaired Data
by: Yacobi, Amitai, et al.
Published: (2025)
by: Yacobi, Amitai, et al.
Published: (2025)
Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models
by: Bergmeister, Andreas, et al.
Published: (2026)
by: Bergmeister, Andreas, et al.
Published: (2026)
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
FodFoM: Fake Outlier Data by Foundation Models Creates Stronger Visual Out-of-Distribution Detector
by: Chen, Jiankang, et al.
Published: (2024)
by: Chen, Jiankang, et al.
Published: (2024)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
The Platonic Representation Hypothesis
by: Huh, Minyoung, et al.
Published: (2024)
by: Huh, Minyoung, et al.
Published: (2024)
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
Adaptive Length Image Tokenization via Recurrent Allocation
by: Duggal, Shivam, et al.
Published: (2024)
by: Duggal, Shivam, et al.
Published: (2024)
Single-pass Adaptive Image Tokenization for Minimum Program Search
by: Duggal, Shivam, et al.
Published: (2025)
by: Duggal, Shivam, et al.
Published: (2025)
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
by: Huh, Minyoung, et al.
Published: (2024)
by: Huh, Minyoung, et al.
Published: (2024)
Tracktention: Leveraging Point Tracking to Attend Videos Faster and Better
by: Lai, Zihang, et al.
Published: (2025)
by: Lai, Zihang, et al.
Published: (2025)
SeaMo: A Season-Aware Multimodal Foundation Model for Remote Sensing
by: Li, Xuyang, et al.
Published: (2024)
by: Li, Xuyang, et al.
Published: (2024)
Make Continual Learning Stronger via C-Flat
by: Bian, Ang, et al.
Published: (2024)
by: Bian, Ang, et al.
Published: (2024)
A Vision Check-up for Language Models
by: Sharma, Pratyusha, et al.
Published: (2024)
by: Sharma, Pratyusha, et al.
Published: (2024)
Unpaired Translation of Point Clouds for Modeling Detector Response
by: Li, Mingyang, et al.
Published: (2025)
by: Li, Mingyang, et al.
Published: (2025)
What Makes for a Good Stereoscopic Image?
by: Tamir, Netanel Y., et al.
Published: (2024)
by: Tamir, Netanel Y., et al.
Published: (2024)
ABC: Achieving Better Control of Multimodal Embeddings using VLMs
by: Schneider, Benjamin, et al.
Published: (2025)
by: Schneider, Benjamin, et al.
Published: (2025)
Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models
by: Lin, Zhiqiu, et al.
Published: (2023)
by: Lin, Zhiqiu, et al.
Published: (2023)
FREE: Faster and Better Data-Free Meta-Learning
by: Wei, Yongxian, et al.
Published: (2024)
by: Wei, Yongxian, et al.
Published: (2024)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
by: Ye, Wenqian, et al.
Published: (2024)
by: Ye, Wenqian, et al.
Published: (2024)
Regularized Distribution Matching Distillation for One-step Unpaired Image-to-Image Translation
by: Rakitin, Denis, et al.
Published: (2024)
by: Rakitin, Denis, et al.
Published: (2024)
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
Attention based End to end network for Offline Writer Identification on Word level data
by: Kumar, Vineet, et al.
Published: (2024)
by: Kumar, Vineet, et al.
Published: (2024)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023)
by: Kim, Jiwon, et al.
Published: (2023)
Foundation Models in Remote Sensing: Evolving from Unimodality to Multimodality
by: Hong, Danfeng, et al.
Published: (2026)
by: Hong, Danfeng, et al.
Published: (2026)
An Information Criterion for Controlled Disentanglement of Multimodal Data
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
Active Learning for Finely-Categorized Image-Text Retrieval by Selecting Hard Negative Unpaired Samples
by: Jo, Dae Ung, et al.
Published: (2024)
by: Jo, Dae Ung, et al.
Published: (2024)
Forget Less by Learning Together through Concept Consolidation
by: Kaushik, Arjun Ramesh, et al.
Published: (2026)
by: Kaushik, Arjun Ramesh, et al.
Published: (2026)
Similar Items
-
In-Context Symmetries: Self-Supervised Learning through Contextual World Models
by: Gupta, Sharut, et al.
Published: (2024) -
Personalized Representation from Personalized Generation
by: Sundaram, Shobhita, et al.
Published: (2024) -
DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data
by: Fu, Stephanie, et al.
Published: (2023) -
Learning Diffusion Models with Flexible Representation Guidance
by: Wang, Chenyu, et al.
Published: (2025) -
When Does Perceptual Alignment Benefit Vision Representations?
by: Sundaram, Shobhita, et al.
Published: (2024)