video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Changli, Li, Yixuan, Yang, Yudong, Zhuang, Jimin, Sun, Guangzhi, Li, Wei, Ma, Zejun, Zhang, Chao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916971080056832
author Tang, Changli
Li, Yixuan
Yang, Yudong
Zhuang, Jimin
Sun, Guangzhi
Li, Wei
Ma, Zejun
Zhang, Chao
author_facet Tang, Changli
Li, Yixuan
Yang, Yudong
Zhuang, Jimin
Sun, Guangzhi
Li, Wei
Ma, Zejun
Zhang, Chao
contents We present video-SALMONN 2, a family of audio-visual large language models that set new state-of-the-art (SOTA) results in video description and question answering (QA). Our core contribution is multi-round direct preference optimisation (MrDPO), paired with a caption-quality objective that jointly rewards completeness and factual accuracy. Unlike standard DPO with a fixed reference policy, MrDPO periodically refreshes the reference by bootstrapping from a newly re-initialised lightweight adapter trained on the latest preferences, avoiding reference staleness and enabling continual improvement. This strategy produces captions that are consistently more detailed and accurate than those from proprietary systems such as GPT-4o and Gemini-1.5 Pro. We further distil these gains by using our model to generate a high-quality video-caption corpus for supervised fine-tuning of new models, transferring benefits beyond captioning to strong performance on complex video-QA tasks. Across widely used audio-visual and visual-only understanding benchmarks (including Video-MME, WorldSense, AVUT, Video-Holmes, DailyOmni, MLVU, and LVBench), our 3B and 7B models achieve SOTA results at comparable scales, while the 72B model surpasses all other open-source systems. Our source code, models, and data are released at \href{https://github.com/bytedance/video-SALMONN-2}{https://github.com/bytedance/video-SALMONN-2}.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15220
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
Tang, Changli
Li, Yixuan
Yang, Yudong
Zhuang, Jimin
Sun, Guangzhi
Li, Wei
Ma, Zejun
Zhang, Chao
Computer Vision and Pattern Recognition
Computation and Language
Sound
We present video-SALMONN 2, a family of audio-visual large language models that set new state-of-the-art (SOTA) results in video description and question answering (QA). Our core contribution is multi-round direct preference optimisation (MrDPO), paired with a caption-quality objective that jointly rewards completeness and factual accuracy. Unlike standard DPO with a fixed reference policy, MrDPO periodically refreshes the reference by bootstrapping from a newly re-initialised lightweight adapter trained on the latest preferences, avoiding reference staleness and enabling continual improvement. This strategy produces captions that are consistently more detailed and accurate than those from proprietary systems such as GPT-4o and Gemini-1.5 Pro. We further distil these gains by using our model to generate a high-quality video-caption corpus for supervised fine-tuning of new models, transferring benefits beyond captioning to strong performance on complex video-QA tasks. Across widely used audio-visual and visual-only understanding benchmarks (including Video-MME, WorldSense, AVUT, Video-Holmes, DailyOmni, MLVU, and LVBench), our 3B and 7B models achieve SOTA results at comparable scales, while the 72B model surpasses all other open-source systems. Our source code, models, and data are released at \href{https://github.com/bytedance/video-SALMONN-2}{https://github.com/bytedance/video-SALMONN-2}.
title video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
Sound
url https://arxiv.org/abs/2506.15220