Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Xiao, Wu, Jianlong, Lin, Zijia, Zhang, Fuzheng, Zhang, Di, Nie, Liqiang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917789578559488
author Wang, Xiao
Wu, Jianlong
Lin, Zijia
Zhang, Fuzheng
Zhang, Di
Nie, Liqiang
author_facet Wang, Xiao
Wu, Jianlong
Lin, Zijia
Zhang, Fuzheng
Zhang, Di
Nie, Liqiang
contents Recently, video-language understanding has achieved great success through large-scale pre-training. However, data scarcity remains a prevailing challenge. This study quantitatively reveals an "impossible trinity" among data quantity, diversity, and quality in pre-training datasets. Recent efforts seek to refine large-scale, diverse ASR datasets compromised by low quality through synthetic annotations. These methods successfully leverage useful information in multimodal video content (frames, tags, ASR transcripts, etc.) to refine the original annotations. Nevertheless, they struggle to mitigate noise within synthetic annotations and lack scalability as the dataset size expands. To address these issues, we introduce the Video DataFlywheel framework, which iteratively refines video annotations with improved noise control methods. For iterative refinement, we first leverage a video-language model to generate synthetic annotations, resulting in a refined dataset. Then, we pre-train on it and fine-tune on human refinement examples for a stronger model. These processes are repeated for continuous improvement. For noise control, we present AdaTaiLr, a novel noise control method that requires weaker assumptions on noise distribution, thereby proving more effective in large datasets with theoretical guarantees. The combination of iterative refinement and AdaTaiLr can achieve better scalability in video-language understanding. Extensive experiments show that our framework outperforms existing data refinement baselines, delivering a 3% performance boost and improving dataset quality with minimal diversity loss. Furthermore, our refined dataset facilitates significant improvements in various video-language understanding tasks, including video question answering and text-video retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19532
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
Wang, Xiao
Wu, Jianlong
Lin, Zijia
Zhang, Fuzheng
Zhang, Di
Nie, Liqiang
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Multimedia
Recently, video-language understanding has achieved great success through large-scale pre-training. However, data scarcity remains a prevailing challenge. This study quantitatively reveals an "impossible trinity" among data quantity, diversity, and quality in pre-training datasets. Recent efforts seek to refine large-scale, diverse ASR datasets compromised by low quality through synthetic annotations. These methods successfully leverage useful information in multimodal video content (frames, tags, ASR transcripts, etc.) to refine the original annotations. Nevertheless, they struggle to mitigate noise within synthetic annotations and lack scalability as the dataset size expands. To address these issues, we introduce the Video DataFlywheel framework, which iteratively refines video annotations with improved noise control methods. For iterative refinement, we first leverage a video-language model to generate synthetic annotations, resulting in a refined dataset. Then, we pre-train on it and fine-tune on human refinement examples for a stronger model. These processes are repeated for continuous improvement. For noise control, we present AdaTaiLr, a novel noise control method that requires weaker assumptions on noise distribution, thereby proving more effective in large datasets with theoretical guarantees. The combination of iterative refinement and AdaTaiLr can achieve better scalability in video-language understanding. Extensive experiments show that our framework outperforms existing data refinement baselines, delivering a 3% performance boost and improving dataset quality with minimal diversity loss. Furthermore, our refined dataset facilitates significant improvements in various video-language understanding tasks, including video question answering and text-video retrieval.
title Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Multimedia
url https://arxiv.org/abs/2409.19532