Multimodal Representation Learning by Alternating Unimodal Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xiaohui, Yoon, Jaehong, Bansal, Mohit, Yao, Huaxiu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
von: Yu, Shoubin, et al.
Veröffentlicht: (2024)
von: Yu, Shoubin, et al.
Veröffentlicht: (2024)
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024)
Continual Learning: Forget-free Winning Subnetworks for Video Representations
von: Kang, Haeyong, et al.
Veröffentlicht: (2023)
von: Kang, Haeyong, et al.
Veröffentlicht: (2023)
Hierarchy-Aware Multimodal Unlearning for Medical AI
von: Wu, Fengli, et al.
Veröffentlicht: (2025)
von: Wu, Fengli, et al.
Veröffentlicht: (2025)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
von: Lin, Han, et al.
Veröffentlicht: (2024)
von: Lin, Han, et al.
Veröffentlicht: (2024)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models
von: Gupta, Sharut, et al.
Veröffentlicht: (2025)
von: Gupta, Sharut, et al.
Veröffentlicht: (2025)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
von: Wei, Yake, et al.
Veröffentlicht: (2024)
von: Wei, Yake, et al.
Veröffentlicht: (2024)
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
von: Hwang, Sunil, et al.
Veröffentlicht: (2022)
von: Hwang, Sunil, et al.
Veröffentlicht: (2022)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
Progressive Fourier Neural Representation for Sequential Video Compilation
von: Kang, Haeyong, et al.
Veröffentlicht: (2023)
von: Kang, Haeyong, et al.
Veröffentlicht: (2023)
MCAQ-YOLO: Morphological Complexity-Aware Quantization for Efficient Object Detection with Curriculum Learning
von: Seo, Yoonjae, et al.
Veröffentlicht: (2025)
von: Seo, Yoonjae, et al.
Veröffentlicht: (2025)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
von: Lee, Jaewoo, et al.
Veröffentlicht: (2023)
von: Lee, Jaewoo, et al.
Veröffentlicht: (2023)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
Deep Regression Representation Learning with Topology
von: Zhang, Shihao, et al.
Veröffentlicht: (2024)
von: Zhang, Shihao, et al.
Veröffentlicht: (2024)
Self-Refining Video Sampling
von: Jang, Sangwon, et al.
Veröffentlicht: (2026)
von: Jang, Sangwon, et al.
Veröffentlicht: (2026)
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
von: Cai, Yichao, et al.
Veröffentlicht: (2025)
von: Cai, Yichao, et al.
Veröffentlicht: (2025)
Beyond Unimodal Learning: The Importance of Integrating Multiple Modalities for Lifelong Learning
von: Sarfraz, Fahad, et al.
Veröffentlicht: (2024)
von: Sarfraz, Fahad, et al.
Veröffentlicht: (2024)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2025)
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2025)
COD: Learning Conditional Invariant Representation for Domain Adaptation Regression
von: Yang, Hao-Ran, et al.
Veröffentlicht: (2024)
von: Yang, Hao-Ran, et al.
Veröffentlicht: (2024)
Dynamic Domain Adaptation-Driven Physics-Informed Graph Representation Learning for AC-OPF
von: Zhu, Hongjie, et al.
Veröffentlicht: (2025)
von: Zhu, Hongjie, et al.
Veröffentlicht: (2025)
Prime Once, then Reprogram Locally: An Efficient Alternative to Black-Box Service Model Adaptation
von: Zhang, Yunbei, et al.
Veröffentlicht: (2026)
von: Zhang, Yunbei, et al.
Veröffentlicht: (2026)
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
von: Li, Po-han, et al.
Veröffentlicht: (2024)
von: Li, Po-han, et al.
Veröffentlicht: (2024)
Principled Multimodal Representation Learning
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
Robust Multimodal Learning with Missing Modalities via Parameter-Efficient Adaptation
von: Reza, Md Kaykobad, et al.
Veröffentlicht: (2023)
von: Reza, Md Kaykobad, et al.
Veröffentlicht: (2023)
Hierarchy-Guided Multimodal Representation Learning for Taxonomic Inference
von: Ahmed, Sk Miraj, et al.
Veröffentlicht: (2026)
von: Ahmed, Sk Miraj, et al.
Veröffentlicht: (2026)
Toward Unified Multimodal Representation Learning for Autonomous Driving
von: Tao, Ximeng, et al.
Veröffentlicht: (2026)
von: Tao, Ximeng, et al.
Veröffentlicht: (2026)
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
von: Zhu, Kangyu, et al.
Veröffentlicht: (2024)
von: Zhu, Kangyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
von: Yoon, Jaehong, et al.
Veröffentlicht: (2024) -
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023) -
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023) -
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
von: Li, Jialu, et al.
Veröffentlicht: (2024) -
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)