On-the-fly Modulation for Balanced Multimodal Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Yake, Hu, Di, Du, Henghui, Wen, Ji-Rong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
Diagnosing and Re-learning for Balanced Multimodal Learning
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
Enhancing multimodal cooperation via sample-level modality valuation
by: Wei, Yake, et al.
Published: (2023)
by: Wei, Yake, et al.
Published: (2023)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
by: Li, Guangyao, et al.
Published: (2024)
by: Li, Guangyao, et al.
Published: (2024)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
by: Lin, Ronghao, et al.
Published: (2024)
by: Lin, Ronghao, et al.
Published: (2024)
An Efficient and Explanatory Image and Text Clustering System with Multimodal Autoencoder Architecture
by: Shi, Tiancheng, et al.
Published: (2024)
by: Shi, Tiancheng, et al.
Published: (2024)
Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
by: Zhang, Haojie, et al.
Published: (2025)
by: Zhang, Haojie, et al.
Published: (2025)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
by: Cai, Dongnuan, et al.
Published: (2026)
by: Cai, Dongnuan, et al.
Published: (2026)
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
by: Wu, Xiaoping, et al.
Published: (2024)
by: Wu, Xiaoping, et al.
Published: (2024)
Multi-layer Learnable Attention Mask for Multimodal Tasks
by: Barrios, Wayner, et al.
Published: (2024)
by: Barrios, Wayner, et al.
Published: (2024)
A Systematic Review on Long-Tailed Learning
by: Zhang, Chongsheng, et al.
Published: (2024)
by: Zhang, Chongsheng, et al.
Published: (2024)
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
by: Barrios, Wayner, et al.
Published: (2025)
by: Barrios, Wayner, et al.
Published: (2025)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
by: Wang, Zhecan, et al.
Published: (2024)
by: Wang, Zhecan, et al.
Published: (2024)
Diffusion Model-Based Video Editing: A Survey
by: Sun, Wenhao, et al.
Published: (2024)
by: Sun, Wenhao, et al.
Published: (2024)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
by: Dong, Hao, et al.
Published: (2026)
by: Dong, Hao, et al.
Published: (2026)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
by: Chen, Weifeng, et al.
Published: (2023)
by: Chen, Weifeng, et al.
Published: (2023)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
by: Li, Liupeng, et al.
Published: (2026)
by: Li, Liupeng, et al.
Published: (2026)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Scaling Spatial Intelligence with Multimodal Foundation Models
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
Multi-level Mixture of Experts for Multimodal Entity Linking
by: Hu, Zhiwei, et al.
Published: (2025)
by: Hu, Zhiwei, et al.
Published: (2025)
Balancing Multimodal Training Through Game-Theoretic Regularization
by: Kontras, Konstantinos, et al.
Published: (2024)
by: Kontras, Konstantinos, et al.
Published: (2024)
OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models
by: Xue, Yida, et al.
Published: (2026)
by: Xue, Yida, et al.
Published: (2026)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
by: Lin, Xiang, et al.
Published: (2025)
by: Lin, Xiang, et al.
Published: (2025)
NVLM: Open Frontier-Class Multimodal LLMs
by: Dai, Wenliang, et al.
Published: (2024)
by: Dai, Wenliang, et al.
Published: (2024)
Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
by: Shen, Meng, et al.
Published: (2024)
by: Shen, Meng, et al.
Published: (2024)
OmniGAIA: Towards Native Omni-Modal AI Agents
by: Li, Xiaoxi, et al.
Published: (2026)
by: Li, Xiaoxi, et al.
Published: (2026)
Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
Machine Unlearning in Hyperbolic vs. Euclidean Multimodal Contrastive Learning: Adapting Alignment Calibration to MERU
by: Vidal, Àlex Pujol, et al.
Published: (2025)
by: Vidal, Àlex Pujol, et al.
Published: (2025)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
by: Lin, Yuanze, et al.
Published: (2025)
by: Lin, Yuanze, et al.
Published: (2025)
Multimodal Real-Time Anomaly Detection and Industrial Applications
by: Verma, Aman, et al.
Published: (2025)
by: Verma, Aman, et al.
Published: (2025)
Reinforcement Learning for Unsupervised Video Summarization with Reward Generator Training
by: Abbasi, Mehryar, et al.
Published: (2024)
by: Abbasi, Mehryar, et al.
Published: (2024)
Continuous Sign Language Recognition System using Deep Learning with MediaPipe Holistic
by: Srivastava, Sharvani, et al.
Published: (2024)
by: Srivastava, Sharvani, et al.
Published: (2024)
Long-tailed Medical Diagnosis with Relation-aware Representation Learning and Iterative Classifier Calibration
by: Pan, Li, et al.
Published: (2025)
by: Pan, Li, et al.
Published: (2025)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
by: Han, Junlin, et al.
Published: (2025)
by: Han, Junlin, et al.
Published: (2025)
HeCoFuse: Cross-Modal Complementary V2X Cooperative Perception with Heterogeneous Sensors
by: Wei, Chuheng, et al.
Published: (2025)
by: Wei, Chuheng, et al.
Published: (2025)
Boosting Facial Action Unit Detection Through Jointly Learning Facial Landmark Detection and Domain Separation and Reconstruction
by: Shang, Ziqiao, et al.
Published: (2023)
by: Shang, Ziqiao, et al.
Published: (2023)
Discover Your Neighbors: Advanced Stable Test-Time Adaptation in Dynamic World
by: Jiang, Qinting, et al.
Published: (2024)
by: Jiang, Qinting, et al.
Published: (2024)
Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
by: Zheng, Shuhong, et al.
Published: (2026)
by: Zheng, Shuhong, et al.
Published: (2026)
Similar Items
-
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024) -
Diagnosing and Re-learning for Balanced Multimodal Learning
by: Wei, Yake, et al.
Published: (2024) -
Enhancing multimodal cooperation via sample-level modality valuation
by: Wei, Yake, et al.
Published: (2023) -
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
by: Li, Guangyao, et al.
Published: (2024) -
End-to-end Semantic-centric Video-based Multimodal Affective Computing
by: Lin, Ronghao, et al.
Published: (2024)