AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Xinlong, Ding, Yue, Lin, Weihong, Hua, Jingyun, Yao, Linli, Shi, Yang, Li, Bozhou, Zhang, Yuanxing, Liu, Qiang, Wan, Pengfei, Wang, Liang, Tan, Tieniu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911205670518784
author Chen, Xinlong
Ding, Yue
Lin, Weihong
Hua, Jingyun
Yao, Linli
Shi, Yang
Li, Bozhou
Zhang, Yuanxing
Liu, Qiang
Wan, Pengfei
Wang, Liang
Tan, Tieniu
author_facet Chen, Xinlong
Ding, Yue
Lin, Weihong
Hua, Jingyun
Yao, Linli
Shi, Yang
Li, Bozhou
Zhang, Yuanxing
Liu, Qiang
Wan, Pengfei
Wang, Liang
Tan, Tieniu
contents Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a powerful audiovisual video captioner driven by the temporal orchestration between audio and visual modalities. We propose a two-stage post-training pipeline: (1) AVoCaDO SFT, which fine-tunes the model on a newly curated dataset of 107K high-quality, temporally-aligned audiovisual captions; and (2) AVoCaDO GRPO, which leverages tailored reward functions to further enhance temporal coherence and dialogue accuracy while regularizing caption length and reducing collapse. Experimental results demonstrate that AVoCaDO significantly outperforms existing open-source models across four audiovisual video captioning benchmarks, and also achieves competitive performance on the VDC and DREAM-1K benchmark under visual-only settings.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10395
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
Chen, Xinlong
Ding, Yue
Lin, Weihong
Hua, Jingyun
Yao, Linli
Shi, Yang
Li, Bozhou
Zhang, Yuanxing
Liu, Qiang
Wan, Pengfei
Wang, Liang
Tan, Tieniu
Computer Vision and Pattern Recognition
Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a powerful audiovisual video captioner driven by the temporal orchestration between audio and visual modalities. We propose a two-stage post-training pipeline: (1) AVoCaDO SFT, which fine-tunes the model on a newly curated dataset of 107K high-quality, temporally-aligned audiovisual captions; and (2) AVoCaDO GRPO, which leverages tailored reward functions to further enhance temporal coherence and dialogue accuracy while regularizing caption length and reducing collapse. Experimental results demonstrate that AVoCaDO significantly outperforms existing open-source models across four audiovisual video captioning benchmarks, and also achieves competitive performance on the VDC and DREAM-1K benchmark under visual-only settings.
title AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.10395