Multimodal Representation Learning by Alternating Unimodal Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xiaohui, Yoon, Jaehong, Bansal, Mohit, Yao, Huaxiu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
by: Yoon, Jaehong, et al.
Published: (2024)
by: Yoon, Jaehong, et al.
Published: (2024)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
by: Zhou, Yiyang, et al.
Published: (2023)
by: Zhou, Yiyang, et al.
Published: (2023)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
by: Sung, Yi-Lin, et al.
Published: (2023)
by: Sung, Yi-Lin, et al.
Published: (2023)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
by: Yu, Shoubin, et al.
Published: (2024)
by: Yu, Shoubin, et al.
Published: (2024)
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
by: Lee, Daeun, et al.
Published: (2024)
by: Lee, Daeun, et al.
Published: (2024)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
by: Yoon, Jaehong, et al.
Published: (2024)
by: Yoon, Jaehong, et al.
Published: (2024)
Continual Learning: Forget-free Winning Subnetworks for Video Representations
by: Kang, Haeyong, et al.
Published: (2023)
by: Kang, Haeyong, et al.
Published: (2023)
Hierarchy-Aware Multimodal Unlearning for Medical AI
by: Wu, Fengli, et al.
Published: (2025)
by: Wu, Fengli, et al.
Published: (2025)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
by: Yu, Shoubin, et al.
Published: (2025)
by: Yu, Shoubin, et al.
Published: (2025)
Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models
by: Gupta, Sharut, et al.
Published: (2025)
by: Gupta, Sharut, et al.
Published: (2025)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
by: Wang, Zun, et al.
Published: (2024)
by: Wang, Zun, et al.
Published: (2024)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
by: Hwang, Sunil, et al.
Published: (2022)
by: Hwang, Sunil, et al.
Published: (2022)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
by: Lee, Daeun, et al.
Published: (2024)
by: Lee, Daeun, et al.
Published: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
Progressive Fourier Neural Representation for Sequential Video Compilation
by: Kang, Haeyong, et al.
Published: (2023)
by: Kang, Haeyong, et al.
Published: (2023)
MCAQ-YOLO: Morphological Complexity-Aware Quantization for Efficient Object Detection with Curriculum Learning
by: Seo, Yoonjae, et al.
Published: (2025)
by: Seo, Yoonjae, et al.
Published: (2025)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023)
by: Lee, Jaewoo, et al.
Published: (2023)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
by: Li, Jialu, et al.
Published: (2025)
by: Li, Jialu, et al.
Published: (2025)
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
by: Yeo, Woongyeong, et al.
Published: (2025)
by: Yeo, Woongyeong, et al.
Published: (2025)
Deep Regression Representation Learning with Topology
by: Zhang, Shihao, et al.
Published: (2024)
by: Zhang, Shihao, et al.
Published: (2024)
Self-Refining Video Sampling
by: Jang, Sangwon, et al.
Published: (2026)
by: Jang, Sangwon, et al.
Published: (2026)
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
by: Cai, Yichao, et al.
Published: (2025)
by: Cai, Yichao, et al.
Published: (2025)
Beyond Unimodal Learning: The Importance of Integrating Multiple Modalities for Lifelong Learning
by: Sarfraz, Fahad, et al.
Published: (2024)
by: Sarfraz, Fahad, et al.
Published: (2024)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
by: Sivakumaran, Nithin, et al.
Published: (2025)
by: Sivakumaran, Nithin, et al.
Published: (2025)
COD: Learning Conditional Invariant Representation for Domain Adaptation Regression
by: Yang, Hao-Ran, et al.
Published: (2024)
by: Yang, Hao-Ran, et al.
Published: (2024)
Dynamic Domain Adaptation-Driven Physics-Informed Graph Representation Learning for AC-OPF
by: Zhu, Hongjie, et al.
Published: (2025)
by: Zhu, Hongjie, et al.
Published: (2025)
Prime Once, then Reprogram Locally: An Efficient Alternative to Black-Box Service Model Adaptation
by: Zhang, Yunbei, et al.
Published: (2026)
by: Zhang, Yunbei, et al.
Published: (2026)
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
Principled Multimodal Representation Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models
by: Lin, Zhiqiu, et al.
Published: (2023)
by: Lin, Zhiqiu, et al.
Published: (2023)
Robust Multimodal Learning with Missing Modalities via Parameter-Efficient Adaptation
by: Reza, Md Kaykobad, et al.
Published: (2023)
by: Reza, Md Kaykobad, et al.
Published: (2023)
Hierarchy-Guided Multimodal Representation Learning for Taxonomic Inference
by: Ahmed, Sk Miraj, et al.
Published: (2026)
by: Ahmed, Sk Miraj, et al.
Published: (2026)
Toward Unified Multimodal Representation Learning for Autonomous Driving
by: Tao, Ximeng, et al.
Published: (2026)
by: Tao, Ximeng, et al.
Published: (2026)
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
by: Zhu, Kangyu, et al.
Published: (2024)
by: Zhu, Kangyu, et al.
Published: (2024)
Similar Items
-
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
by: Yoon, Jaehong, et al.
Published: (2024) -
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
by: Zhou, Yiyang, et al.
Published: (2023) -
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
by: Sung, Yi-Lin, et al.
Published: (2023) -
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
by: Li, Jialu, et al.
Published: (2024) -
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
by: Wang, Xiyao, et al.
Published: (2024)