AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Jinda, Li, Jinghan, Gao, Yuan, Wu, Junkang, Wu, Jiancan, Wang, Xiang, He, Xiangnan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
by: Lu, Jinda, et al.
Published: (2026)
by: Lu, Jinda, et al.
Published: (2026)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
by: Lu, Jinda, et al.
Published: (2025)
by: Lu, Jinda, et al.
Published: (2025)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
RePO: Understanding Preference Learning Through ReLU-Based Optimization
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
by: Ji, Wence, et al.
Published: (2025)
by: Ji, Wence, et al.
Published: (2025)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Unified Parameter-Efficient Unlearning for LLMs
by: Ding, Chenlu, et al.
Published: (2024)
by: Ding, Chenlu, et al.
Published: (2024)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
by: Li, Jinghan, et al.
Published: (2026)
by: Li, Jinghan, et al.
Published: (2026)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
by: Huang, Kexin, et al.
Published: (2025)
by: Huang, Kexin, et al.
Published: (2025)
Robust Preference Optimization via Dynamic Target Margins
by: Sun, Jie, et al.
Published: (2025)
by: Sun, Jie, et al.
Published: (2025)
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
by: Lu, Jinda, et al.
Published: (2026)
by: Lu, Jinda, et al.
Published: (2026)
On Negative-aware Preference Optimization for Recommendation
by: Ding, Chenlu, et al.
Published: (2025)
by: Ding, Chenlu, et al.
Published: (2025)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
by: Huang, Kexin, et al.
Published: (2026)
by: Huang, Kexin, et al.
Published: (2026)
RosePO: Aligning LLM-based Recommenders with Human Values
by: Liao, Jiayi, et al.
Published: (2024)
by: Liao, Jiayi, et al.
Published: (2024)
Causal-HalBench: Uncovering LVLMs Object Hallucinations Through Causal Intervention
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
Direct Multi-Turn Preference Optimization for Language Agents
by: Shi, Wentao, et al.
Published: (2024)
by: Shi, Wentao, et al.
Published: (2024)
Reinforced Prompt Personalization for Recommendation with Large Language Models
by: Mao, Wenyu, et al.
Published: (2024)
by: Mao, Wenyu, et al.
Published: (2024)
R^2-Mem: Reflective Experience for Memory Search
by: Wang, Xinyuan, et al.
Published: (2026)
by: Wang, Xinyuan, et al.
Published: (2026)
Align, Don't Divide: Revisiting the LoRA Architecture in Multi-Task Learning
by: Liu, Jinda, et al.
Published: (2025)
by: Liu, Jinda, et al.
Published: (2025)
LLaRA: Large Language-Recommendation Assistant
by: Liao, Jiayi, et al.
Published: (2023)
by: Liao, Jiayi, et al.
Published: (2023)
Addressing Missing Data Issue for Diffusion-based Recommendation
by: Mao, Wenyu, et al.
Published: (2025)
by: Mao, Wenyu, et al.
Published: (2025)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
by: Ding, Chenlu, et al.
Published: (2025)
by: Ding, Chenlu, et al.
Published: (2025)
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
by: Wu, Linquan, et al.
Published: (2026)
by: Wu, Linquan, et al.
Published: (2026)
Aligning Multimodal LLM with Human Preference: A Survey
by: Yu, Tao, et al.
Published: (2025)
by: Yu, Tao, et al.
Published: (2025)
Leave No Patient Behind: Enhancing Medication Recommendation for Rare Disease Patients
by: Zhao, Zihao, et al.
Published: (2024)
by: Zhao, Zihao, et al.
Published: (2024)
Customizing Language Models with Instance-wise LoRA for Sequential Recommendation
by: Kong, Xiaoyu, et al.
Published: (2024)
by: Kong, Xiaoyu, et al.
Published: (2024)
Boosting Few-Shot Learning via Attentive Feature Regularization
by: Zhu, Xingyu, et al.
Published: (2024)
by: Zhu, Xingyu, et al.
Published: (2024)
DiffGAD: A Diffusion-based Unsupervised Graph Anomaly Detector
by: Li, Jinghan, et al.
Published: (2024)
by: Li, Jinghan, et al.
Published: (2024)
R-LoRA: Randomized Multi-Head LoRA for Efficient Multi-Task Learning
by: Liu, Jinda, et al.
Published: (2025)
by: Liu, Jinda, et al.
Published: (2025)
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
by: Corrado, Nicholas E., et al.
Published: (2025)
by: Corrado, Nicholas E., et al.
Published: (2025)
Lower-Left Partial AUC: An Effective and Efficient Optimization Metric for Recommendation
by: Shi, Wentao, et al.
Published: (2024)
by: Shi, Wentao, et al.
Published: (2024)
Aligning LLMs with Individual Preferences via Interaction
by: Wu, Shujin, et al.
Published: (2024)
by: Wu, Shujin, et al.
Published: (2024)
Aligning CodeLLMs with Direct Preference Optimization
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
Adaptive Self-supervised Robust Clustering for Unstructured Data with Unknown Cluster Number
by: Ding, Chen-Lu, et al.
Published: (2024)
by: Ding, Chen-Lu, et al.
Published: (2024)
MiniOneRec: An Open-Source Framework for Scaling Generative Recommendation
by: Kong, Xiaoyu, et al.
Published: (2025)
by: Kong, Xiaoyu, et al.
Published: (2025)
TiAda: A Time-scale Adaptive Algorithm for Nonconvex Minimax Optimization
by: Li, Xiang, et al.
Published: (2022)
by: Li, Xiang, et al.
Published: (2022)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
ViPO: Visual Preference Optimization at Scale
by: Li, Ming, et al.
Published: (2026)
by: Li, Ming, et al.
Published: (2026)
Similar Items
-
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
by: Lu, Jinda, et al.
Published: (2026) -
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
by: Lu, Jinda, et al.
Published: (2025) -
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024) -
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024) -
RePO: Understanding Preference Learning Through ReLU-Based Optimization
by: Wu, Junkang, et al.
Published: (2025)