Distilling Vision-Language Models on Millions of Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910411640537088 |
|---|---|
| author | Zhao, Yue Zhao, Long Zhou, Xingyi Wu, Jialin Chu, Chun-Te Miao, Hui Schroff, Florian Adam, Hartwig Liu, Ting Gong, Boqing Krähenbühl, Philipp Yuan, Liangzhe |
| author_facet | Zhao, Yue Zhao, Long Zhou, Xingyi Wu, Jialin Chu, Chun-Te Miao, Hui Schroff, Florian Adam, Hartwig Liu, Ting Gong, Boqing Krähenbühl, Philipp Yuan, Liangzhe |
| contents | The recent advance in vision-language models is largely attributed to the abundance of image-text data. We aim to replicate this success for video-language models, but there simply is not enough human-curated video-text data available. We thus resort to fine-tuning a video-language model from a strong image-language baseline with synthesized instructional data. The resulting video model by video-instruction-tuning (VIIT) is then used to auto-label millions of videos to generate high-quality captions. We show the adapted video-language model performs well on a wide range of video-language benchmarks. For instance, it surpasses the best prior result on open-ended NExT-QA by 2.8%. Besides, our model generates detailed descriptions for previously unseen videos, which provide better textual supervision than existing methods. Experiments show that a video-language dual-encoder model contrastively trained on these auto-generated captions is 3.8% better than the strongest baseline that also leverages vision-language models. Our best model outperforms state-of-the-art methods on MSR-VTT zero-shot text-to-video retrieval by 6%. As a side product, we generate the largest video caption dataset to date. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_06129 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Distilling Vision-Language Models on Millions of Videos Zhao, Yue Zhao, Long Zhou, Xingyi Wu, Jialin Chu, Chun-Te Miao, Hui Schroff, Florian Adam, Hartwig Liu, Ting Gong, Boqing Krähenbühl, Philipp Yuan, Liangzhe Computer Vision and Pattern Recognition The recent advance in vision-language models is largely attributed to the abundance of image-text data. We aim to replicate this success for video-language models, but there simply is not enough human-curated video-text data available. We thus resort to fine-tuning a video-language model from a strong image-language baseline with synthesized instructional data. The resulting video model by video-instruction-tuning (VIIT) is then used to auto-label millions of videos to generate high-quality captions. We show the adapted video-language model performs well on a wide range of video-language benchmarks. For instance, it surpasses the best prior result on open-ended NExT-QA by 2.8%. Besides, our model generates detailed descriptions for previously unseen videos, which provide better textual supervision than existing methods. Experiments show that a video-language dual-encoder model contrastively trained on these auto-generated captions is 3.8% better than the strongest baseline that also leverages vision-language models. Our best model outperforms state-of-the-art methods on MSR-VTT zero-shot text-to-video retrieval by 6%. As a side product, we generate the largest video caption dataset to date. |
| title | Distilling Vision-Language Models on Millions of Videos |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2401.06129 |