Can Text-to-Video Generation help Video-Language Alignment?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zanella, Luca, Mancini, Massimiliano, Menapace, Willi, Tulyakov, Sergey, Wang, Yiming, Ricci, Elisa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917966915829760
author Zanella, Luca
Mancini, Massimiliano
Menapace, Willi
Tulyakov, Sergey
Wang, Yiming
Ricci, Elisa
author_facet Zanella, Luca
Mancini, Massimiliano
Menapace, Willi
Tulyakov, Sergey
Wang, Yiming
Ricci, Elisa
contents Recent video-language alignment models are trained on sets of videos, each with an associated positive caption and a negative caption generated by large language models. A problem with this procedure is that negative captions may introduce linguistic biases, i.e., concepts are seen only as negatives and never associated with a video. While a solution would be to collect videos for the negative captions, existing databases lack the fine-grained variations needed to cover all possible negatives. In this work, we study whether synthetic videos can help to overcome this issue. Our preliminary analysis with multiple generators shows that, while promising on some tasks, synthetic videos harm the performance of the model on others. We hypothesize this issue is linked to noise (semantic and visual) in the generated videos and develop a method, SynViTA, that accounts for those. SynViTA dynamically weights the contribution of each synthetic video based on how similar its target caption is w.r.t. the real counterpart. Moreover, a semantic consistency loss makes the model focus on fine-grained differences across captions, rather than differences in video appearance. Experiments show that, on average, SynViTA improves over existing methods on VideoCon test sets and SSv2-Temporal, SSv2-Events, and ATP-Hard benchmarks, being a first promising step for using synthetic videos when learning video-language models.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Text-to-Video Generation help Video-Language Alignment?
Zanella, Luca
Mancini, Massimiliano
Menapace, Willi
Tulyakov, Sergey
Wang, Yiming
Ricci, Elisa
Computer Vision and Pattern Recognition
Recent video-language alignment models are trained on sets of videos, each with an associated positive caption and a negative caption generated by large language models. A problem with this procedure is that negative captions may introduce linguistic biases, i.e., concepts are seen only as negatives and never associated with a video. While a solution would be to collect videos for the negative captions, existing databases lack the fine-grained variations needed to cover all possible negatives. In this work, we study whether synthetic videos can help to overcome this issue. Our preliminary analysis with multiple generators shows that, while promising on some tasks, synthetic videos harm the performance of the model on others. We hypothesize this issue is linked to noise (semantic and visual) in the generated videos and develop a method, SynViTA, that accounts for those. SynViTA dynamically weights the contribution of each synthetic video based on how similar its target caption is w.r.t. the real counterpart. Moreover, a semantic consistency loss makes the model focus on fine-grained differences across captions, rather than differences in video appearance. Experiments show that, on average, SynViTA improves over existing methods on VideoCon test sets and SSv2-Temporal, SSv2-Events, and ATP-Hard benchmarks, being a first promising step for using synthetic videos when learning video-language models.
title Can Text-to-Video Generation help Video-Language Alignment?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18507