Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zitian, Liao, Yue, Rong, Kang, Rao, Fengyun, Yang, Yibo, Liu, Si |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024)
by: Zhang, Mengxi, et al.
Published: (2024)
MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection
by: Wang, Zitian, et al.
Published: (2024)
by: Wang, Zitian, et al.
Published: (2024)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
by: Zhou, Zitang, et al.
Published: (2025)
by: Zhou, Zitang, et al.
Published: (2025)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
by: Yue, Xinli, et al.
Published: (2025)
by: Yue, Xinli, et al.
Published: (2025)
MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling
by: Yang, Jian, et al.
Published: (2024)
by: Yang, Jian, et al.
Published: (2024)
Enhancing 3D Lane Detection and Topology Reasoning with 2D Lane Priors
by: Li, Han, et al.
Published: (2024)
by: Li, Han, et al.
Published: (2024)
SAIL: Self-Amplified Iterative Learning for Diffusion Model Alignment with Minimal Human Feedback
by: He, Xiaoxuan, et al.
Published: (2026)
by: He, Xiaoxuan, et al.
Published: (2026)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
by: Hao, Yunzhuo, et al.
Published: (2025)
by: Hao, Yunzhuo, et al.
Published: (2025)
SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs
by: Lou, Haoran, et al.
Published: (2026)
by: Lou, Haoran, et al.
Published: (2026)
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
by: Su, Yongyi, et al.
Published: (2025)
by: Su, Yongyi, et al.
Published: (2025)
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
by: Li, Mingxiao, et al.
Published: (2025)
by: Li, Mingxiao, et al.
Published: (2025)
Multi-Modal Generative Embedding Model
by: Ma, Feipeng, et al.
Published: (2024)
by: Ma, Feipeng, et al.
Published: (2024)
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
by: Jiang, Songtao, et al.
Published: (2024)
by: Jiang, Songtao, et al.
Published: (2024)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
Quantifying and Enhancing Multi-modal Robustness with Modality Preference
by: Yang, Zequn, et al.
Published: (2024)
by: Yang, Zequn, et al.
Published: (2024)
MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization
by: Li, Yong, et al.
Published: (2026)
by: Li, Yong, et al.
Published: (2026)
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment
by: Suo, Yucheng, et al.
Published: (2025)
by: Suo, Yucheng, et al.
Published: (2025)
FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs
by: Hao, Jing, et al.
Published: (2024)
by: Hao, Jing, et al.
Published: (2024)
Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth
by: Chen, Mingrui, et al.
Published: (2026)
by: Chen, Mingrui, et al.
Published: (2026)
Preference Alignment on Diffusion Model: A Comprehensive Survey for Image Generation and Editing
by: Wu, Sihao, et al.
Published: (2025)
by: Wu, Sihao, et al.
Published: (2025)
The Geometry of Compromise: Unlocking Generative Capabilities via Controllable Modality Alignment
by: Liu, Hongyuan, et al.
Published: (2026)
by: Liu, Hongyuan, et al.
Published: (2026)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
by: Li, Xudong, et al.
Published: (2025)
by: Li, Xudong, et al.
Published: (2025)
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
by: Zhu, Muzhi, et al.
Published: (2025)
by: Zhu, Muzhi, et al.
Published: (2025)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
Learning Modality Knowledge Alignment for Cross-Modality Transfer
by: Ma, Wenxuan, et al.
Published: (2024)
by: Ma, Wenxuan, et al.
Published: (2024)
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
by: Zhang, Kejia, et al.
Published: (2025)
by: Zhang, Kejia, et al.
Published: (2025)
SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs
by: Wang, Jiahui, et al.
Published: (2025)
by: Wang, Jiahui, et al.
Published: (2025)
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
by: Zhang, Guanghao, et al.
Published: (2025)
by: Zhang, Guanghao, et al.
Published: (2025)
Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
by: Zhao, Pengfei, et al.
Published: (2025)
by: Zhao, Pengfei, et al.
Published: (2025)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
by: Sun, Peiwen, et al.
Published: (2026)
by: Sun, Peiwen, et al.
Published: (2026)
Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs
by: Si, Guangzong, et al.
Published: (2025)
by: Si, Guangzong, et al.
Published: (2025)
Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations
by: Cui, Yibo, et al.
Published: (2025)
by: Cui, Yibo, et al.
Published: (2025)
FDCT: Frequency-Aware Decomposition and Cross-Modal Token-Alignment for Multi-Sensor Target Classification
by: Sami, Shoaib Meraj, et al.
Published: (2025)
by: Sami, Shoaib Meraj, et al.
Published: (2025)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
by: Liu, Yangzhou, et al.
Published: (2024)
by: Liu, Yangzhou, et al.
Published: (2024)
Similar Items
-
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
by: Yang, Jie, et al.
Published: (2025) -
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024) -
MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection
by: Wang, Zitian, et al.
Published: (2024) -
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
by: Zhao, Xiangyu, et al.
Published: (2025) -
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
by: Zhou, Zitang, et al.
Published: (2025)