Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Songtao, Zhang, Yan, Chen, Ruizhe, Hu, Tianxiang, Jin, Yeying, He, Qinglin, Feng, Yang, Wu, Jian, Liu, Zuozhu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
by: Jiang, Songtao, et al.
Published: (2025)
by: Jiang, Songtao, et al.
Published: (2025)
Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models
by: Jiang, Songtao, et al.
Published: (2024)
by: Jiang, Songtao, et al.
Published: (2024)
Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
by: Zhao, Pengfei, et al.
Published: (2025)
by: Zhao, Pengfei, et al.
Published: (2025)
UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making
by: Jiang, Songtao, et al.
Published: (2025)
by: Jiang, Songtao, et al.
Published: (2025)
Implicit Preference Alignment for Human Image Animation
by: Wang, Yuanzhi, et al.
Published: (2026)
by: Wang, Yuanzhi, et al.
Published: (2026)
PIP-MM: Pre-Integrating Prompt Information into Visual Encoding via Existing MLLM Structures
by: Wu, Tianxiang, et al.
Published: (2024)
by: Wu, Tianxiang, et al.
Published: (2024)
InstructX: Towards Unified Visual Editing with MLLM Guidance
by: Mou, Chong, et al.
Published: (2025)
by: Mou, Chong, et al.
Published: (2025)
IOSVLM: A 3D Vision-Language Model for Unified Dental Diagnosis from Intraoral Scans
by: Xiong, Huimin, et al.
Published: (2026)
by: Xiong, Huimin, et al.
Published: (2026)
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
by: Cai, Weitong, et al.
Published: (2024)
by: Cai, Weitong, et al.
Published: (2024)
Fair-MoE: Fairness-Oriented Mixture of Experts in Vision-Language Models
by: Wang, Peiran, et al.
Published: (2025)
by: Wang, Peiran, et al.
Published: (2025)
Towards Distribution-Agnostic Generalized Category Discovery
by: Bai, Jianhong, et al.
Published: (2023)
by: Bai, Jianhong, et al.
Published: (2023)
SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models
by: Lian, Jiesong, et al.
Published: (2026)
by: Lian, Jiesong, et al.
Published: (2026)
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
by: Ran, Ran, et al.
Published: (2026)
by: Ran, Ran, et al.
Published: (2026)
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024)
by: Wijaya, Robert, et al.
Published: (2024)
FAIntbench: A Holistic and Precise Benchmark for Bias Evaluation in Text-to-Image Models
by: Luo, Hanjun, et al.
Published: (2024)
by: Luo, Hanjun, et al.
Published: (2024)
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
by: Xin, Yi, et al.
Published: (2025)
by: Xin, Yi, et al.
Published: (2025)
PX2Tooth: Reconstructing the 3D Point Cloud Teeth from a Single Panoramic X-ray
by: Ma, Wen, et al.
Published: (2024)
by: Ma, Wen, et al.
Published: (2024)
Multimodal Alignment and Fusion: A Survey
by: Li, Songtao, et al.
Published: (2024)
by: Li, Songtao, et al.
Published: (2024)
KPL: Training-Free Medical Knowledge Mining of Vision-Language Models
by: Liu, Jiaxiang, et al.
Published: (2025)
by: Liu, Jiaxiang, et al.
Published: (2025)
Decoupling Dark Knowledge via Block-wise Logit Distillation for Feature-level Alignment
by: Yu, Chengting, et al.
Published: (2024)
by: Yu, Chengting, et al.
Published: (2024)
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
by: Wang, Zitian, et al.
Published: (2025)
by: Wang, Zitian, et al.
Published: (2025)
MedThink: Explaining Medical Visual Question Answering via Multimodal Decision-Making Rationale
by: Gai, Xiaotang, et al.
Published: (2024)
by: Gai, Xiaotang, et al.
Published: (2024)
MLLM-as-a-Judge Exhibits Model Preference Bias
by: Koyama, Shuitsu, et al.
Published: (2026)
by: Koyama, Shuitsu, et al.
Published: (2026)
Leveraging MLLM Embeddings and Attribute Smoothing for Compositional Zero-Shot Learning
by: Yan, Xudong, et al.
Published: (2024)
by: Yan, Xudong, et al.
Published: (2024)
Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs
by: Cao, Rui, et al.
Published: (2024)
by: Cao, Rui, et al.
Published: (2024)
Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models
by: Xia, Yuexuan, et al.
Published: (2025)
by: Xia, Yuexuan, et al.
Published: (2025)
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
by: Jiang, Songtao, et al.
Published: (2025)
by: Jiang, Songtao, et al.
Published: (2025)
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment
by: Jiang, Lifan, et al.
Published: (2025)
by: Jiang, Lifan, et al.
Published: (2025)
SHAPE : Self-Improved Visual Preference Alignment by Iteratively Generating Holistic Winner
by: Chen, Kejia, et al.
Published: (2025)
by: Chen, Kejia, et al.
Published: (2025)
Trustworthy and Fair SkinGPT-R1 for Democratizing Dermatological Reasoning across Diverse Ethnicities
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos
by: Jiang, Songtao, et al.
Published: (2026)
by: Jiang, Songtao, et al.
Published: (2026)
Robustness-Guided Image Synthesis for Data-Free Quantization
by: Bai, Jianhong, et al.
Published: (2023)
by: Bai, Jianhong, et al.
Published: (2023)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
by: Yan, Ziang, et al.
Published: (2024)
by: Yan, Ziang, et al.
Published: (2024)
BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM
by: Wen, Haiquan, et al.
Published: (2025)
by: Wen, Haiquan, et al.
Published: (2025)
Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset
by: Deng, Ziye, et al.
Published: (2025)
by: Deng, Ziye, et al.
Published: (2025)
Decoupling Bias, Aligning Distributions: Synergistic Fairness Optimization for Deepfake Detection
by: Ding, Feng, et al.
Published: (2025)
by: Ding, Feng, et al.
Published: (2025)
Quantifying and Enhancing Multi-modal Robustness with Modality Preference
by: Yang, Zequn, et al.
Published: (2024)
by: Yang, Zequn, et al.
Published: (2024)
Raindrop Clarity: A Dual-Focused Dataset for Day and Night Raindrop Removal
by: Jin, Yeying, et al.
Published: (2024)
by: Jin, Yeying, et al.
Published: (2024)
Bridging the Gap: Multi-Level Cross-Modality Joint Alignment for Visible-Infrared Person Re-Identification
by: Liang, Tengfei, et al.
Published: (2023)
by: Liang, Tengfei, et al.
Published: (2023)
Similar Items
-
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
by: Jiang, Songtao, et al.
Published: (2025) -
Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models
by: Jiang, Songtao, et al.
Published: (2024) -
Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
by: Zhao, Pengfei, et al.
Published: (2025) -
UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
by: Zhang, Wei, et al.
Published: (2025) -
CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making
by: Jiang, Songtao, et al.
Published: (2025)