Temporal Preference Optimization for Long-Form Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Rui, Wang, Xiaohan, Zhang, Yuhui, Zohar, Orr, Wang, Zeyu, Yeung-Levy, Serena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909761796046848
author Li, Rui
Wang, Xiaohan
Zhang, Yuhui
Zohar, Orr
Wang, Zeyu
Yeung-Levy, Serena
author_facet Li, Rui
Wang, Xiaohan
Zhang, Yuhui
Zohar, Orr
Wang, Zeyu
Yeung-Levy, Serena
contents Despite significant advancements in video large multimodal models (video-LMMs), achieving effective temporal grounding in long-form videos remains a challenge for existing models. To address this limitation, we propose Temporal Preference Optimization (TPO), a novel post-training framework designed to enhance the temporal grounding capabilities of video-LMMs through preference learning. TPO adopts a self-training approach that enables models to differentiate between well-grounded and less accurate temporal responses by leveraging curated preference datasets at two granularities: localized temporal grounding, which focuses on specific video segments, and comprehensive temporal grounding, which captures extended temporal dependencies across entire video sequences. By optimizing on these preference datasets, TPO significantly enhances temporal understanding while reducing reliance on manually annotated data. Extensive experiments on three long-form video understanding benchmarks--LongVideoBench, MLVU, and Video-MME--demonstrate the effectiveness of TPO across two state-of-the-art video-LMMs. Notably, LLaVA-Video-TPO establishes itself as the leading 7B model on the Video-MME benchmark, underscoring the potential of TPO as a scalable and efficient solution for advancing temporal reasoning in long-form video understanding. Project page: https://ruili33.github.io/tpo_website.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13919
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Temporal Preference Optimization for Long-Form Video Understanding
Li, Rui
Wang, Xiaohan
Zhang, Yuhui
Zohar, Orr
Wang, Zeyu
Yeung-Levy, Serena
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Robotics
Despite significant advancements in video large multimodal models (video-LMMs), achieving effective temporal grounding in long-form videos remains a challenge for existing models. To address this limitation, we propose Temporal Preference Optimization (TPO), a novel post-training framework designed to enhance the temporal grounding capabilities of video-LMMs through preference learning. TPO adopts a self-training approach that enables models to differentiate between well-grounded and less accurate temporal responses by leveraging curated preference datasets at two granularities: localized temporal grounding, which focuses on specific video segments, and comprehensive temporal grounding, which captures extended temporal dependencies across entire video sequences. By optimizing on these preference datasets, TPO significantly enhances temporal understanding while reducing reliance on manually annotated data. Extensive experiments on three long-form video understanding benchmarks--LongVideoBench, MLVU, and Video-MME--demonstrate the effectiveness of TPO across two state-of-the-art video-LMMs. Notably, LLaVA-Video-TPO establishes itself as the leading 7B model on the Video-MME benchmark, underscoring the potential of TPO as a scalable and efficient solution for advancing temporal reasoning in long-form video understanding. Project page: https://ruili33.github.io/tpo_website.
title Temporal Preference Optimization for Long-Form Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Robotics
url https://arxiv.org/abs/2501.13919