Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Ziyang, Xu, Ruiyang, Xing, Zhenghao, Chu, Yunfei, Wang, Yuxuan, He, Jinzheng, Xu, Jin, Heng, Pheng-Ann, Yu, Kai, Lin, Junyang, Chng, Eng Siong, Chen, Xie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911519041650688
author Ma, Ziyang
Xu, Ruiyang
Xing, Zhenghao
Chu, Yunfei
Wang, Yuxuan
He, Jinzheng
Xu, Jin
Heng, Pheng-Ann
Yu, Kai
Lin, Junyang
Chng, Eng Siong
Chen, Xie
author_facet Ma, Ziyang
Xu, Ruiyang
Xing, Zhenghao
Chu, Yunfei
Wang, Yuxuan
He, Jinzheng
Xu, Jin
Heng, Pheng-Ann
Yu, Kai
Lin, Junyang
Chng, Eng Siong
Chen, Xie
contents Fine-grained perception of multimodal information is critical for advancing human-AI interaction. With recent progress in audio-visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel, have emerged as a promising paradigm for achieving richer understanding and reasoning. However, their capacity to capture and describe fine-grained details remains limited explored. In this work, we present a systematic and comprehensive investigation of omni detailed perception from the perspectives of the data pipeline, models, and benchmark. We first identify an inherent "co-growth" between detail and hallucination in current OLMs. To address this, we propose Omni-Detective, an agentic data generation pipeline integrating tool-calling, to autonomously produce highly detailed yet minimally hallucinatory multimodal data. Based on the data generated with Omni-Detective, we train two captioning models: Audio-Captioner for audio-only detailed perception, and Omni-Captioner for audio-visual detailed perception. Under the cascade evaluation protocol, Audio-Captioner achieves the best performance on MMAU and MMAR among all open-source models, surpassing Gemini 2.5 Flash and delivering performance comparable to Gemini 2.5 Pro. On existing detailed captioning benchmarks, Omni-Captioner sets a new state-of-the-art on VDC and achieves the best trade-off between detail and hallucination on the video-SALMONN 2 testset. Given the absence of a dedicated benchmark for omni detailed perception, we design Omni-Cloze, a novel cloze-style evaluation for detailed audio, visual, and audio-visual captioning that ensures stable, efficient, and reliable assessment. Experimental results and analysis demonstrate the effectiveness of Omni-Detective in generating high-quality detailed captions, as well as the superiority of Omni-Cloze in evaluating such detailed captions.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12720
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
Ma, Ziyang
Xu, Ruiyang
Xing, Zhenghao
Chu, Yunfei
Wang, Yuxuan
He, Jinzheng
Xu, Jin
Heng, Pheng-Ann
Yu, Kai
Lin, Junyang
Chng, Eng Siong
Chen, Xie
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Sound
Fine-grained perception of multimodal information is critical for advancing human-AI interaction. With recent progress in audio-visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel, have emerged as a promising paradigm for achieving richer understanding and reasoning. However, their capacity to capture and describe fine-grained details remains limited explored. In this work, we present a systematic and comprehensive investigation of omni detailed perception from the perspectives of the data pipeline, models, and benchmark. We first identify an inherent "co-growth" between detail and hallucination in current OLMs. To address this, we propose Omni-Detective, an agentic data generation pipeline integrating tool-calling, to autonomously produce highly detailed yet minimally hallucinatory multimodal data. Based on the data generated with Omni-Detective, we train two captioning models: Audio-Captioner for audio-only detailed perception, and Omni-Captioner for audio-visual detailed perception. Under the cascade evaluation protocol, Audio-Captioner achieves the best performance on MMAU and MMAR among all open-source models, surpassing Gemini 2.5 Flash and delivering performance comparable to Gemini 2.5 Pro. On existing detailed captioning benchmarks, Omni-Captioner sets a new state-of-the-art on VDC and achieves the best trade-off between detail and hallucination on the video-SALMONN 2 testset. Given the absence of a dedicated benchmark for omni detailed perception, we design Omni-Cloze, a novel cloze-style evaluation for detailed audio, visual, and audio-visual captioning that ensures stable, efficient, and reliable assessment. Experimental results and analysis demonstrate the effectiveness of Omni-Detective in generating high-quality detailed captions, as well as the superiority of Omni-Cloze in evaluating such detailed captions.
title Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
topic Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Sound
url https://arxiv.org/abs/2510.12720