VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Dahun, Piergiovanni, AJ, Mallya, Ganesh, Angelova, Anelia
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908310928621568
author Kim, Dahun
Piergiovanni, AJ
Mallya, Ganesh
Angelova, Anelia
author_facet Kim, Dahun
Piergiovanni, AJ
Mallya, Ganesh
Angelova, Anelia
contents We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing benchmarks focused on static image-text compositionality or isolated single-event videos, our benchmark targets alignment in continuous multi-event videos. Leveraging video-text datasets with temporally localized event captions (e.g. ActivityNet-Captions, YouCook2), we construct two compositional benchmarks, ActivityNet-Comp and YouCook2-Comp. We create challenging negative samples with subtle temporal disruptions such as reordering, action word replacement, partial captioning, and combined disruptions. These benchmarks comprehensively test models' compositional sensitivity across extended, cohesive video-text sequences. To improve model performance, we propose a hierarchical pairwise preference loss that strengthens alignment with temporally accurate pairs and gradually penalizes increasingly disrupted ones, encouraging fine-grained compositional learning. To mitigate the limited availability of densely annotated video data, we introduce a pretraining strategy that concatenates short video-caption pairs to simulate multi-event sequences. We evaluate video-text foundational models and large multimodal models (LMMs) on our benchmark, identifying both strengths and areas for improvement in compositionality. Overall, our work provides a comprehensive framework for evaluating and enhancing model capabilities in achieving fine-grained, temporally coherent video-text alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2504_03970
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
Kim, Dahun
Piergiovanni, AJ
Mallya, Ganesh
Angelova, Anelia
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Information Retrieval
We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing benchmarks focused on static image-text compositionality or isolated single-event videos, our benchmark targets alignment in continuous multi-event videos. Leveraging video-text datasets with temporally localized event captions (e.g. ActivityNet-Captions, YouCook2), we construct two compositional benchmarks, ActivityNet-Comp and YouCook2-Comp. We create challenging negative samples with subtle temporal disruptions such as reordering, action word replacement, partial captioning, and combined disruptions. These benchmarks comprehensively test models' compositional sensitivity across extended, cohesive video-text sequences. To improve model performance, we propose a hierarchical pairwise preference loss that strengthens alignment with temporally accurate pairs and gradually penalizes increasingly disrupted ones, encouraging fine-grained compositional learning. To mitigate the limited availability of densely annotated video data, we introduce a pretraining strategy that concatenates short video-caption pairs to simulate multi-event sequences. We evaluate video-text foundational models and large multimodal models (LMMs) on our benchmark, identifying both strengths and areas for improvement in compositionality. Overall, our work provides a comprehensive framework for evaluating and enhancing model capabilities in achieving fine-grained, temporally coherent video-text alignment.
title VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2504.03970