Understanding Alignment in Multimodal LLMs: A Comprehensive Study
Fuente:
arXiv
Guardado en:
| Autores principales: | Amirloo, Elmira, Fauconnier, Jean-Philippe, Roesmann, Christoph, Kerl, Christian, Boney, Rinu, Qian, Yusu, Wang, Zirui, Dehghan, Afshin, Yang, Yinfei, Gan, Zhe, Grasch, Peter |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
por: Qian, Yusu, et al.
Publicado: (2024)
por: Qian, Yusu, et al.
Publicado: (2024)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
por: Qian, Yusu, et al.
Publicado: (2024)
por: Qian, Yusu, et al.
Publicado: (2024)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
por: Daxberger, Erik, et al.
Publicado: (2025)
por: Daxberger, Erik, et al.
Publicado: (2025)
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
por: Ye, Shaokai, et al.
Publicado: (2026)
por: Ye, Shaokai, et al.
Publicado: (2026)
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
por: Zhang, Haotian, et al.
Publicado: (2024)
por: Zhang, Haotian, et al.
Publicado: (2024)
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
por: Feng, Di, et al.
Publicado: (2025)
por: Feng, Di, et al.
Publicado: (2025)
DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
por: Narayan, Kartik, et al.
Publicado: (2025)
por: Narayan, Kartik, et al.
Publicado: (2025)
PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
por: Qian, Yusu, et al.
Publicado: (2025)
por: Qian, Yusu, et al.
Publicado: (2025)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
por: You, Keen, et al.
Publicado: (2024)
por: You, Keen, et al.
Publicado: (2024)
UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
por: Fu, Tsu-Jui, et al.
Publicado: (2025)
por: Fu, Tsu-Jui, et al.
Publicado: (2025)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
por: Tian, Rui, et al.
Publicado: (2025)
por: Tian, Rui, et al.
Publicado: (2025)
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
por: Bachmann, Roman, et al.
Publicado: (2025)
por: Bachmann, Roman, et al.
Publicado: (2025)
UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning
por: Tian, Rui, et al.
Publicado: (2025)
por: Tian, Rui, et al.
Publicado: (2025)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
por: Xu, Mingze, et al.
Publicado: (2025)
por: Xu, Mingze, et al.
Publicado: (2025)
Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
por: Qian, Yusu, et al.
Publicado: (2025)
por: Qian, Yusu, et al.
Publicado: (2025)
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
por: Qian, Yusu, et al.
Publicado: (2025)
por: Qian, Yusu, et al.
Publicado: (2025)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
por: Jaiswal, Ajay, et al.
Publicado: (2023)
por: Jaiswal, Ajay, et al.
Publicado: (2023)
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
por: McKinzie, Brandon, et al.
Publicado: (2024)
por: McKinzie, Brandon, et al.
Publicado: (2024)
Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping
por: Lazarow, Justin, et al.
Publicado: (2025)
por: Lazarow, Justin, et al.
Publicado: (2025)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
por: Lai, Zhengfeng, et al.
Publicado: (2024)
por: Lai, Zhengfeng, et al.
Publicado: (2024)
AToken: A Unified Tokenizer for Vision
por: Lu, Jiasen, et al.
Publicado: (2025)
por: Lu, Jiasen, et al.
Publicado: (2025)
Guiding Instruction-based Image Editing via Multimodal Large Language Models
por: Fu, Tsu-Jui, et al.
Publicado: (2023)
por: Fu, Tsu-Jui, et al.
Publicado: (2023)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
por: Xu, Mingze, et al.
Publicado: (2024)
por: Xu, Mingze, et al.
Publicado: (2024)
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
por: Kou, Qian, et al.
Publicado: (2026)
por: Kou, Qian, et al.
Publicado: (2026)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
por: Ye, Hanrong, et al.
Publicado: (2024)
por: Ye, Hanrong, et al.
Publicado: (2024)
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
por: Wang, Zirui, et al.
Publicado: (2024)
por: Wang, Zirui, et al.
Publicado: (2024)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
por: Li, Zhangheng, et al.
Publicado: (2024)
por: Li, Zhangheng, et al.
Publicado: (2024)
Cubify Anything: Scaling Indoor 3D Object Detection
por: Lazarow, Justin, et al.
Publicado: (2024)
por: Lazarow, Justin, et al.
Publicado: (2024)
Reliable Viscosity Calculation from High-Pressure Equilibrium Molecular Dynamics: Case Study of 2,2,4-Trimethylhexane
por: Toraman, Gözdenur, et al.
Publicado: (2026)
por: Toraman, Gözdenur, et al.
Publicado: (2026)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
por: Shukor, Mustafa, et al.
Publicado: (2024)
por: Shukor, Mustafa, et al.
Publicado: (2024)
Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs
por: Zheng, Huan, et al.
Publicado: (2026)
por: Zheng, Huan, et al.
Publicado: (2026)
STable AutoCorrelation Integral Estimator (STACIE): Robust and accurate transport properties from molecular dynamics simulations
por: Toraman, Gözdenur, et al.
Publicado: (2025)
por: Toraman, Gözdenur, et al.
Publicado: (2025)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
por: Mousavi, Pooneh, et al.
Publicado: (2025)
por: Mousavi, Pooneh, et al.
Publicado: (2025)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
por: An, Wenbin, et al.
Publicado: (2025)
por: An, Wenbin, et al.
Publicado: (2025)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
por: Patel, Alkesh, et al.
Publicado: (2025)
por: Patel, Alkesh, et al.
Publicado: (2025)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
por: Fu, Chaoyou, et al.
Publicado: (2024)
por: Fu, Chaoyou, et al.
Publicado: (2024)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
por: Song, Tingyu, et al.
Publicado: (2025)
por: Song, Tingyu, et al.
Publicado: (2025)
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation
por: Wang, Yongxin, et al.
Publicado: (2024)
por: Wang, Yongxin, et al.
Publicado: (2024)
Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
Understanding LLMs: A Comprehensive Overview from Training to Inference
por: Liu, Yiheng, et al.
Publicado: (2024)
por: Liu, Yiheng, et al.
Publicado: (2024)
Ejemplares similares
-
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
por: Qian, Yusu, et al.
Publicado: (2024) -
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
por: Qian, Yusu, et al.
Publicado: (2024) -
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
por: Daxberger, Erik, et al.
Publicado: (2025) -
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
por: Ye, Shaokai, et al.
Publicado: (2026) -
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
por: Zhang, Haotian, et al.
Publicado: (2024)