AVC-DPO: Aligned Video Captioning via Direct Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Jiyang, Li, Hengyi, Du, Yifan, Zhao, Wayne Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916821504884736
author Tang, Jiyang
Li, Hengyi
Du, Yifan
Zhao, Wayne Xin
author_facet Tang, Jiyang
Li, Hengyi
Du, Yifan
Zhao, Wayne Xin
contents Although video multimodal large language models (video MLLMs) have achieved substantial progress in video captioning tasks, it remains challenging to adjust the focal emphasis of video captions according to human preferences. To address this limitation, we propose Aligned Video Captioning via Direct Preference Optimization (AVC-DPO), a post-training framework designed to enhance captioning capabilities in video MLLMs through preference alignment. Our approach designs enhanced prompts that specifically target temporal dynamics and spatial information-two key factors that humans care about when watching a video-thereby incorporating human-centric preferences. AVC-DPO leverages the same foundation model's caption generation responses under varied prompt conditions to conduct preference-aware training and caption alignment. Using this framework, we have achieved exceptional performance in the LOVE@CVPR'25 Workshop Track 1A: Video Detailed Captioning Challenge, achieving first place on the Video Detailed Captioning (VDC) benchmark according to the VDCSCORE evaluation metric.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
Tang, Jiyang
Li, Hengyi
Du, Yifan
Zhao, Wayne Xin
Computer Vision and Pattern Recognition
Although video multimodal large language models (video MLLMs) have achieved substantial progress in video captioning tasks, it remains challenging to adjust the focal emphasis of video captions according to human preferences. To address this limitation, we propose Aligned Video Captioning via Direct Preference Optimization (AVC-DPO), a post-training framework designed to enhance captioning capabilities in video MLLMs through preference alignment. Our approach designs enhanced prompts that specifically target temporal dynamics and spatial information-two key factors that humans care about when watching a video-thereby incorporating human-centric preferences. AVC-DPO leverages the same foundation model's caption generation responses under varied prompt conditions to conduct preference-aware training and caption alignment. Using this framework, we have achieved exceptional performance in the LOVE@CVPR'25 Workshop Track 1A: Video Detailed Captioning Challenge, achieving first place on the Video Detailed Captioning (VDC) benchmark according to the VDCSCORE evaluation metric.
title AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.01492