Toward Robust Multimodal Learning using Multimodal Foundational Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xianbing, Poria, Soujanya, Li, Xuejiao, Chen, Yixin, Tang, Buzhou |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
by: Chia, Yew Ken, et al.
Published: (2024)
by: Chia, Yew Ken, et al.
Published: (2024)
The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles
by: Toh, Vernon Y. H., et al.
Published: (2025)
by: Toh, Vernon Y. H., et al.
Published: (2025)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
by: Sun, Qi, et al.
Published: (2024)
by: Sun, Qi, et al.
Published: (2024)
Towards Robust Instruction Tuning on Multimodal Large Language Models
by: Han, Wei, et al.
Published: (2024)
by: Han, Wei, et al.
Published: (2024)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
by: Han, Wei, et al.
Published: (2023)
by: Han, Wei, et al.
Published: (2023)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
by: Ghosal, Deepanway, et al.
Published: (2024)
by: Ghosal, Deepanway, et al.
Published: (2024)
Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned
by: Ong, Brandon, et al.
Published: (2025)
by: Ong, Brandon, et al.
Published: (2025)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
by: Li, Wenbin, et al.
Published: (2026)
by: Li, Wenbin, et al.
Published: (2026)
Robust Multimodal Large Language Models Against Modality Conflict
by: Zhang, Zongmeng, et al.
Published: (2025)
by: Zhang, Zongmeng, et al.
Published: (2025)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
by: Shangguan, Ziyao, et al.
Published: (2024)
by: Shangguan, Ziyao, et al.
Published: (2024)
Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
by: Zhang, Zhiqian, et al.
Published: (2026)
by: Zhang, Zhiqian, et al.
Published: (2026)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
by: LASA Team, et al.
Published: (2025)
by: LASA Team, et al.
Published: (2025)
LFTR: Learning-Free Token Reduction for Multimodal Large Language Models
by: Zhao, Zihui, et al.
Published: (2025)
by: Zhao, Zihui, et al.
Published: (2025)
Many-Shot In-Context Learning in Multimodal Foundation Models
by: Jiang, Yixing, et al.
Published: (2024)
by: Jiang, Yixing, et al.
Published: (2024)
Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
by: Buckley, Thomas, et al.
Published: (2023)
by: Buckley, Thomas, et al.
Published: (2023)
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study
by: Wang, Chenguang, et al.
Published: (2024)
by: Wang, Chenguang, et al.
Published: (2024)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Graphic Design with Large Multimodal Model
by: Cheng, Yutao, et al.
Published: (2024)
by: Cheng, Yutao, et al.
Published: (2024)
MLLM-CL: Continual Learning for Multimodal Large Language Models
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
Multimodal Chain-of-Thought Reasoning in Language Models
by: Zhang, Zhuosheng, et al.
Published: (2023)
by: Zhang, Zhuosheng, et al.
Published: (2023)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
by: Huang, Jinsheng, et al.
Published: (2024)
by: Huang, Jinsheng, et al.
Published: (2024)
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
by: Suglia, Alessandro, et al.
Published: (2024)
by: Suglia, Alessandro, et al.
Published: (2024)
MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models
by: Paik, Gio, et al.
Published: (2025)
by: Paik, Gio, et al.
Published: (2025)
HEMM: Holistic Evaluation of Multimodal Foundation Models
by: Liang, Paul Pu, et al.
Published: (2024)
by: Liang, Paul Pu, et al.
Published: (2024)
Keyword-Oriented Multimodal Modeling for Euphemism Identification
by: Hu, Yuxue, et al.
Published: (2025)
by: Hu, Yuxue, et al.
Published: (2025)
VITA: Towards Open-Source Interactive Omni Multimodal LLM
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
Model Composition for Multimodal Large Language Models
by: Chen, Chi, et al.
Published: (2024)
by: Chen, Chi, et al.
Published: (2024)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
by: Wang, Yuqing, et al.
Published: (2023)
by: Wang, Yuqing, et al.
Published: (2023)
Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations
by: Li, Yanshu
Published: (2025)
by: Li, Yanshu
Published: (2025)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
by: Wu, Chenyuan, et al.
Published: (2025)
by: Wu, Chenyuan, et al.
Published: (2025)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
Towards Understanding Graphical Perception in Large Multimodal Models
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Bridging the Gap Between Multimodal Foundation Models and World Models
by: He, Xuehai
Published: (2025)
by: He, Xuehai
Published: (2025)
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
by: Ashraf, Tajamul, et al.
Published: (2025)
by: Ashraf, Tajamul, et al.
Published: (2025)
Jailbreaking Multimodal Large Language Models using Multi-Clip Video
by: Kang, Choongwon, et al.
Published: (2026)
by: Kang, Choongwon, et al.
Published: (2026)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
by: Zhao, Haozhe, et al.
Published: (2025)
by: Zhao, Haozhe, et al.
Published: (2025)
Towards Visual Text Grounding of Multimodal Large Language Model
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Large Multimodal Agents: A Survey
by: Xie, Junlin, et al.
Published: (2024)
by: Xie, Junlin, et al.
Published: (2024)
Similar Items
-
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
by: Chia, Yew Ken, et al.
Published: (2024) -
The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles
by: Toh, Vernon Y. H., et al.
Published: (2025) -
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
by: Sun, Qi, et al.
Published: (2024) -
Towards Robust Instruction Tuning on Multimodal Large Language Models
by: Han, Wei, et al.
Published: (2024) -
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
by: Han, Wei, et al.
Published: (2023)