Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Rui, Jiang, Yuming, Schlichtkrull, Michael, Vlachos, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024)
by: Wijaya, Robert, et al.
Published: (2024)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
by: Han, Su Ho, et al.
Published: (2025)
by: Han, Su Ho, et al.
Published: (2025)
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024)
by: Zhang, Mengxi, et al.
Published: (2024)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
by: Jiang, Songtao, et al.
Published: (2024)
by: Jiang, Songtao, et al.
Published: (2024)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
Preference Score Distillation: Leveraging 2D Rewards to Align Text-to-3D Generation with Human Preference
by: Leng, Jiaqi, et al.
Published: (2026)
by: Leng, Jiaqi, et al.
Published: (2026)
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
by: Li, Xudong, et al.
Published: (2025)
by: Li, Xudong, et al.
Published: (2025)
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
by: Ishikawa, Reina, et al.
Published: (2025)
by: Ishikawa, Reina, et al.
Published: (2025)
Visual Jigsaw Post-Training Improves MLLMs
by: Wu, Penghao, et al.
Published: (2025)
by: Wu, Penghao, et al.
Published: (2025)
KIND: Knowledge Integration and Diversion for Training Decomposable Models
by: Xie, Yucheng, et al.
Published: (2024)
by: Xie, Yucheng, et al.
Published: (2024)
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
by: Lu, Chaochao, et al.
Published: (2024)
by: Lu, Chaochao, et al.
Published: (2024)
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
by: Zhang, Kejia, et al.
Published: (2025)
by: Zhang, Kejia, et al.
Published: (2025)
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
by: Wang, Zitian, et al.
Published: (2025)
by: Wang, Zitian, et al.
Published: (2025)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
by: Yu, Tianyu, et al.
Published: (2023)
by: Yu, Tianyu, et al.
Published: (2023)
Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation
by: Kim, Jungeun, et al.
Published: (2024)
by: Kim, Jungeun, et al.
Published: (2024)
Awaker2.5-VL: Stably Scaling MLLMs with Parameter-Efficient Mixture of Experts
by: Long, Jinqiang, et al.
Published: (2024)
by: Long, Jinqiang, et al.
Published: (2024)
SEF-MAP: Subspace-Decomposed Expert Fusion for Robust Multimodal HD Map Prediction
by: Fu, Haoxiang, et al.
Published: (2026)
by: Fu, Haoxiang, et al.
Published: (2026)
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
AME: Aligned Manifold Entropy for Robust Vision-Language Distillation
by: Cao, Guiming, et al.
Published: (2025)
by: Cao, Guiming, et al.
Published: (2025)
DeferredSeg: A Multi-Expert Deferral Framework for Trustworthy Medical Image Segmentation
by: Tian, Qiuyu, et al.
Published: (2026)
by: Tian, Qiuyu, et al.
Published: (2026)
Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time Adaptation
by: Zhang, Rongyu, et al.
Published: (2024)
by: Zhang, Rongyu, et al.
Published: (2024)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
by: Zhang, Bob, et al.
Published: (2025)
by: Zhang, Bob, et al.
Published: (2025)
Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
by: Ma, Xueqi, et al.
Published: (2025)
by: Ma, Xueqi, et al.
Published: (2025)
Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
by: Li, Yuanshuai, et al.
Published: (2025)
by: Li, Yuanshuai, et al.
Published: (2025)
Shadow Generation with Decomposed Mask Prediction and Attentive Shadow Filling
by: Tao, Xinhao, et al.
Published: (2023)
by: Tao, Xinhao, et al.
Published: (2023)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025)
by: Kim, Sanghwan, et al.
Published: (2025)
FDIO: Frequency Decomposed Inertial Odometry
by: Zhang, Shanshan, et al.
Published: (2025)
by: Zhang, Shanshan, et al.
Published: (2025)
Does Visual Token Pruning Improve Calibration? An Empirical Study on Confidence in MLLMs
by: Tan, Kaizhen
Published: (2026)
by: Tan, Kaizhen
Published: (2026)
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
by: Zhang, Gengyuan, et al.
Published: (2025)
by: Zhang, Gengyuan, et al.
Published: (2025)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
Chimera: Improving Generalist Model with Domain-Specific Experts
by: Peng, Tianshuo, et al.
Published: (2024)
by: Peng, Tianshuo, et al.
Published: (2024)
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
by: Xu, Xiao, et al.
Published: (2025)
by: Xu, Xiao, et al.
Published: (2025)
Improving Progressive Generation with Decomposable Flow Matching
by: Haji-Ali, Moayed, et al.
Published: (2025)
by: Haji-Ali, Moayed, et al.
Published: (2025)
Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and Conditioning
by: Ruan, Penghui, et al.
Published: (2024)
by: Ruan, Penghui, et al.
Published: (2024)
Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion Models
by: Korkmaz, Cansu, et al.
Published: (2025)
by: Korkmaz, Cansu, et al.
Published: (2025)
Similar Items
-
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024) -
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
by: Han, Su Ho, et al.
Published: (2025) -
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024) -
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024) -
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
by: Jiang, Songtao, et al.
Published: (2024)