Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Maniparambil, Mayug, Akshulakov, Raiymbek, Djilali, Yasser Abdelaziz Dahou, Narayan, Sanath, Singh, Ankit, O'Connor, Noel E. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Vision and Language Encoders Represent the World Similarly?
by: Maniparambil, Mayug, et al.
Published: (2024)
by: Maniparambil, Mayug, et al.
Published: (2024)
ViSpeR: Multilingual Audio-Visual Speech Recognition
by: Narayan, Sanath, et al.
Published: (2024)
by: Narayan, Sanath, et al.
Published: (2024)
Hold-One-Shot-Out (HOSO) for Validation-Free Few-Shot CLIP Adapters
by: Vorster, Chris, et al.
Published: (2026)
by: Vorster, Chris, et al.
Published: (2026)
Underrepresented in Foundation Model Pretraining Data? A One-Shot Probe
by: Vorster, Chris, et al.
Published: (2026)
by: Vorster, Chris, et al.
Published: (2026)
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
by: Hoehing, Nils, et al.
Published: (2025)
by: Hoehing, Nils, et al.
Published: (2025)
Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction?
by: Hashmi, Anam, et al.
Published: (2026)
by: Hashmi, Anam, et al.
Published: (2026)
Ensemble Learning with Sparse Hypercolumns
by: Dietlmeier, Julia, et al.
Published: (2026)
by: Dietlmeier, Julia, et al.
Published: (2026)
Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation
by: Sirotkin, Kirill, et al.
Published: (2024)
by: Sirotkin, Kirill, et al.
Published: (2024)
Vision-Language Models Can't See the Obvious
by: Dahou, Yasser, et al.
Published: (2025)
by: Dahou, Yasser, et al.
Published: (2025)
Test-Time Adaptation with SaLIP: A Cascade of SAM and CLIP for Zero shot Medical Image Segmentation
by: Aleem, Sidra, et al.
Published: (2024)
by: Aleem, Sidra, et al.
Published: (2024)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
Falcon Perception
by: Bevli, Aviraj, et al.
Published: (2026)
by: Bevli, Aviraj, et al.
Published: (2026)
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
Falcon2-11B Technical Report
by: Malartic, Quentin, et al.
Published: (2024)
by: Malartic, Quentin, et al.
Published: (2024)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
Towards Label-Free Brain Tumor Segmentation: Unsupervised Learning with Multimodal MRI
by: Comas-Quiles, Gerard, et al.
Published: (2025)
by: Comas-Quiles, Gerard, et al.
Published: (2025)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
by: Gupta, Akshita, et al.
Published: (2024)
by: Gupta, Akshita, et al.
Published: (2024)
Semantic Alignment of Unimodal Medical Text and Vision Representations
by: Di Folco, Maxime, et al.
Published: (2025)
by: Di Folco, Maxime, et al.
Published: (2025)
Assessing and Learning Alignment of Unimodal Vision and Language Models
by: Zhang, Le, et al.
Published: (2024)
by: Zhang, Le, et al.
Published: (2024)
Parameter-Free Bio-Inspired Channel Attention for Enhanced Cardiac MRI Reconstruction
by: Hashmi, Anam, et al.
Published: (2025)
by: Hashmi, Anam, et al.
Published: (2025)
Accelerating Cardiac MRI Reconstruction with CMRatt: An Attention-Driven Approach
by: Hashmi, Anam, et al.
Published: (2024)
by: Hashmi, Anam, et al.
Published: (2024)
Multimodal Representation Learning by Alternating Unimodal Adaptation
by: Zhang, Xiaohui, et al.
Published: (2023)
by: Zhang, Xiaohui, et al.
Published: (2023)
Batch Augmentation with Unimodal Fine-tuning for Multimodal Learning
by: Kabir, H M Dipu, et al.
Published: (2025)
by: Kabir, H M Dipu, et al.
Published: (2025)
VLSM-Ensemble: Ensembling CLIP-based Vision-Language Models for Enhanced Medical Image Segmentation
by: Dietlmeier, Julia, et al.
Published: (2025)
by: Dietlmeier, Julia, et al.
Published: (2025)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
by: Cavagnero, Niccolò, et al.
Published: (2026)
by: Cavagnero, Niccolò, et al.
Published: (2026)
Foundation Models in Remote Sensing: Evolving from Unimodality to Multimodality
by: Hong, Danfeng, et al.
Published: (2026)
by: Hong, Danfeng, et al.
Published: (2026)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
by: Li, Zhuowan, et al.
Published: (2022)
by: Li, Zhuowan, et al.
Published: (2022)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
An accurate detection is not all you need to combat label noise in web-noisy datasets
by: Albert, Paul, et al.
Published: (2024)
by: Albert, Paul, et al.
Published: (2024)
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
by: Zhu, Jing, et al.
Published: (2025)
by: Zhu, Jing, et al.
Published: (2025)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
by: Pang, Ziqi, et al.
Published: (2023)
by: Pang, Ziqi, et al.
Published: (2023)
Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning
by: Li, Jianxiong, et al.
Published: (2024)
by: Li, Jianxiong, et al.
Published: (2024)
HyperTopo-Adapters: Geometry- and Topology-Aware Segmentation of Leaf Lesions on Frozen Encoders
by: Ndubuisi, Chimdi Walter, et al.
Published: (2025)
by: Ndubuisi, Chimdi Walter, et al.
Published: (2025)
DINOv3 as a Frozen Encoder for CRPS-Oriented Probabilistic Rainfall Nowcasting
by: Filho, Luciano Araujo Dourado, et al.
Published: (2025)
by: Filho, Luciano Araujo Dourado, et al.
Published: (2025)
Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models
by: Gupta, Sharut, et al.
Published: (2025)
by: Gupta, Sharut, et al.
Published: (2025)
Reinforcement Learning meets Masked Video Modeling : Trajectory-Guided Adaptive Token Selection
by: Rai, Ayush K., et al.
Published: (2025)
by: Rai, Ayush K., et al.
Published: (2025)
Cross-Modal Knowledge Distillation for PET-Free Amyloid-Beta Detection from MRI
by: Chiumento, Francesco, et al.
Published: (2026)
by: Chiumento, Francesco, et al.
Published: (2026)
Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification
by: Khan, Md Ashik, et al.
Published: (2025)
by: Khan, Md Ashik, et al.
Published: (2025)
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
Similar Items
-
Do Vision and Language Encoders Represent the World Similarly?
by: Maniparambil, Mayug, et al.
Published: (2024) -
ViSpeR: Multilingual Audio-Visual Speech Recognition
by: Narayan, Sanath, et al.
Published: (2024) -
Hold-One-Shot-Out (HOSO) for Validation-Free Few-Shot CLIP Adapters
by: Vorster, Chris, et al.
Published: (2026) -
Underrepresented in Foundation Model Pretraining Data? A One-Shot Probe
by: Vorster, Chris, et al.
Published: (2026) -
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
by: Hoehing, Nils, et al.
Published: (2025)