VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Haojian, Chen, Haodong, Wu, Shengqiong, Luo, Meng, Fu, Jinlan, Du, Xinya, Zhang, Hanwang, Fei, Hao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908324672307200
author Huang, Haojian
Chen, Haodong
Wu, Shengqiong
Luo, Meng
Fu, Jinlan
Du, Xinya
Zhang, Hanwang
Fei, Hao
author_facet Huang, Haojian
Chen, Haodong
Wu, Shengqiong
Luo, Meng
Fu, Jinlan
Du, Xinya
Zhang, Hanwang
Fei, Hao
contents Large Video Models (LVMs) built upon Large Language Models (LLMs) have shown promise in video understanding but often suffer from misalignment with human intuition and video hallucination issues. To address these challenges, we introduce VistaDPO, a novel framework for Video Hierarchical Spatial-Temporal Direct Preference Optimization. VistaDPO enhances text-video preference alignment across three hierarchical levels: i) Instance Level, aligning overall video content with responses; ii) Temporal Level, aligning video temporal semantics with event descriptions; and iii) Perceptive Level, aligning spatial objects with language tokens. Given the lack of datasets for fine-grained video-language preference alignment, we construct VistaDPO-7k, a dataset of 7.2K QA pairs annotated with chosen and rejected responses, along with spatial-temporal grounding information such as timestamps, keyframes, and bounding boxes. Extensive experiments on benchmarks such as Video Hallucination, Video QA, and Captioning performance tasks demonstrate that VistaDPO significantly improves the performance of existing LVMs, effectively mitigating video-language misalignment and hallucination. The code and data are available at https://github.com/HaroldChen19/VistaDPO.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13122
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
Huang, Haojian
Chen, Haodong
Wu, Shengqiong
Luo, Meng
Fu, Jinlan
Du, Xinya
Zhang, Hanwang
Fei, Hao
Computer Vision and Pattern Recognition
Machine Learning
Large Video Models (LVMs) built upon Large Language Models (LLMs) have shown promise in video understanding but often suffer from misalignment with human intuition and video hallucination issues. To address these challenges, we introduce VistaDPO, a novel framework for Video Hierarchical Spatial-Temporal Direct Preference Optimization. VistaDPO enhances text-video preference alignment across three hierarchical levels: i) Instance Level, aligning overall video content with responses; ii) Temporal Level, aligning video temporal semantics with event descriptions; and iii) Perceptive Level, aligning spatial objects with language tokens. Given the lack of datasets for fine-grained video-language preference alignment, we construct VistaDPO-7k, a dataset of 7.2K QA pairs annotated with chosen and rejected responses, along with spatial-temporal grounding information such as timestamps, keyframes, and bounding boxes. Extensive experiments on benchmarks such as Video Hallucination, Video QA, and Captioning performance tasks demonstrate that VistaDPO significantly improves the performance of existing LVMs, effectively mitigating video-language misalignment and hallucination. The code and data are available at https://github.com/HaroldChen19/VistaDPO.
title VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2504.13122