Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Sarkar, Pritam, Etemad, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models
by: Sarkar, Pritam, et al.
Published: (2025)
by: Sarkar, Pritam, et al.
Published: (2025)
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
by: Roy, Shuvendu, et al.
Published: (2025)
by: Roy, Shuvendu, et al.
Published: (2025)
Consistency-guided Prompt Learning for Vision-Language Models
by: Roy, Shuvendu, et al.
Published: (2023)
by: Roy, Shuvendu, et al.
Published: (2023)
Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
by: Sarkar, Pritam, et al.
Published: (2024)
by: Sarkar, Pritam, et al.
Published: (2024)
CycleCrash: A Dataset of Bicycle Collision Videos for Collision Prediction and Analysis
by: Desai, Nishq Poorav, et al.
Published: (2024)
by: Desai, Nishq Poorav, et al.
Published: (2024)
CollideNet: Hierarchical Multi-scale Video Representation Learning with Disentanglement for Time-To-Collision Forecasting
by: Desai, Nishq Poorav, et al.
Published: (2026)
by: Desai, Nishq Poorav, et al.
Published: (2026)
Impact of Strategic Sampling and Supervision Policies on Semi-supervised Learning
by: Roy, Shuvendu, et al.
Published: (2022)
by: Roy, Shuvendu, et al.
Published: (2022)
Exploring the Boundaries of Semi-Supervised Facial Expression Recognition using In-Distribution, Out-of-Distribution, and Unconstrained Data
by: Roy, Shuvendu, et al.
Published: (2023)
by: Roy, Shuvendu, et al.
Published: (2023)
Partial Label Learning for Emotion Recognition from EEG
by: Zhang, Guangyi, et al.
Published: (2023)
by: Zhang, Guangyi, et al.
Published: (2023)
Some Optimizers are More Equal: Understanding the Role of Optimizers in Group Fairness
by: Kolahdouzi, Mojtaba, et al.
Published: (2025)
by: Kolahdouzi, Mojtaba, et al.
Published: (2025)
Diffusion Models with Deterministic Normalizing Flow Priors
by: Zand, Mohsen, et al.
Published: (2023)
by: Zand, Mohsen, et al.
Published: (2023)
Unmasking Deepfakes: Masked Autoencoding Spatiotemporal Transformers for Enhanced Video Forgery Detection
by: Das, Sayantan, et al.
Published: (2023)
by: Das, Sayantan, et al.
Published: (2023)
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
by: Zhang, Ruohong, et al.
Published: (2024)
by: Zhang, Ruohong, et al.
Published: (2024)
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
by: Mishra, Pritam, et al.
Published: (2025)
by: Mishra, Pritam, et al.
Published: (2025)
Scaling Up Semi-supervised Learning with Unconstrained Unlabelled Data
by: Roy, Shuvendu, et al.
Published: (2023)
by: Roy, Shuvendu, et al.
Published: (2023)
Human Pose Estimation from Ambiguous Pressure Recordings with Spatio-temporal Masked Transformers
by: Davoodnia, Vandad, et al.
Published: (2023)
by: Davoodnia, Vandad, et al.
Published: (2023)
Consistency-Guided Asynchronous Contrastive Tuning for Few-Shot Class-Incremental Tuning of Foundation Models
by: Roy, Shuvendu, et al.
Published: (2024)
by: Roy, Shuvendu, et al.
Published: (2024)
Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
by: Patnaik, Sohan, et al.
Published: (2025)
by: Patnaik, Sohan, et al.
Published: (2025)
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
by: Baid, Ami, et al.
Published: (2026)
by: Baid, Ami, et al.
Published: (2026)
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
by: Kim, Dain, et al.
Published: (2026)
by: Kim, Dain, et al.
Published: (2026)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
by: Salamatian, Ali, et al.
Published: (2026)
by: Salamatian, Ali, et al.
Published: (2026)
Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization
by: Qiu, Xinyu, et al.
Published: (2026)
by: Qiu, Xinyu, et al.
Published: (2026)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
by: Yan, Ziang, et al.
Published: (2024)
by: Yan, Ziang, et al.
Published: (2024)
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding
by: Pereira, Joao, et al.
Published: (2025)
by: Pereira, Joao, et al.
Published: (2025)
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
by: Jeon, Byungwoo, et al.
Published: (2026)
by: Jeon, Byungwoo, et al.
Published: (2026)
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
by: Huang, Haojian, et al.
Published: (2025)
by: Huang, Haojian, et al.
Published: (2025)
Self-Refining Video Sampling
by: Jang, Sangwon, et al.
Published: (2026)
by: Jang, Sangwon, et al.
Published: (2026)
LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization
by: Huang, Zhenpeng, et al.
Published: (2026)
by: Huang, Zhenpeng, et al.
Published: (2026)
Debiasing Multimodal Large Language Models via Noise-Aware Preference Optimization
by: Zhang, Zefeng, et al.
Published: (2025)
by: Zhang, Zefeng, et al.
Published: (2025)
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
by: Luo, Jiayun, et al.
Published: (2024)
by: Luo, Jiayun, et al.
Published: (2024)
SkelFormer: Markerless 3D Pose and Shape Estimation using Skeletal Transformers
by: Davoodnia, Vandad, et al.
Published: (2024)
by: Davoodnia, Vandad, et al.
Published: (2024)
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
by: Yashima, Daichi, et al.
Published: (2026)
by: Yashima, Daichi, et al.
Published: (2026)
Self-Supervised Human Activity Recognition with Localized Time-Frequency Contrastive Representation Learning
by: Taghanaki, Setareh Rahimi, et al.
Published: (2022)
by: Taghanaki, Setareh Rahimi, et al.
Published: (2022)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
by: Mishra, Pritam, et al.
Published: (2026)
by: Mishra, Pritam, et al.
Published: (2026)
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
by: Tran, Tuyen, et al.
Published: (2025)
by: Tran, Tuyen, et al.
Published: (2025)
MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization
by: Pappa, Massimiliano, et al.
Published: (2024)
by: Pappa, Massimiliano, et al.
Published: (2024)
Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation
by: Du, Jie, et al.
Published: (2025)
by: Du, Jie, et al.
Published: (2025)
Similar Items
-
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models
by: Sarkar, Pritam, et al.
Published: (2025) -
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
by: Roy, Shuvendu, et al.
Published: (2025) -
Consistency-guided Prompt Learning for Vision-Language Models
by: Roy, Shuvendu, et al.
Published: (2023) -
Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
by: Sarkar, Pritam, et al.
Published: (2024) -
CycleCrash: A Dataset of Bicycle Collision Videos for Collision Prediction and Analysis
by: Desai, Nishq Poorav, et al.
Published: (2024)