MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lei, Zhenxin, Gao, Zhangwei, Tian, Changyao, Cui, Erfei, Chen, Guanzhou, Yang, Danni, Duan, Yuchen, Wang, Zhaokai, Li, Wenhao, Wang, Weiyun, Zhao, Xiangyu, Ji, Jiayi, Qiao, Yu, Wang, Wenhai, Luo, Gen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909850075660288
author Lei, Zhenxin
Gao, Zhangwei
Tian, Changyao
Cui, Erfei
Chen, Guanzhou
Yang, Danni
Duan, Yuchen
Wang, Zhaokai
Li, Wenhao
Wang, Weiyun
Zhao, Xiangyu
Ji, Jiayi
Qiao, Yu
Wang, Wenhai
Luo, Gen
author_facet Lei, Zhenxin
Gao, Zhangwei
Tian, Changyao
Cui, Erfei
Chen, Guanzhou
Yang, Danni
Duan, Yuchen
Wang, Zhaokai
Li, Wenhao
Wang, Weiyun
Zhao, Xiangyu
Ji, Jiayi
Qiao, Yu
Wang, Wenhai
Luo, Gen
contents Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains. In this task, current open-source models present a large performance gap with commercial ones, which limits various applications such as data synthesis. To bridge the gap, this paper proposes CapFlow, a novel multi-agent collaboration workflow. CapFlow demonstrates for the first time that, by capitalizing on open-source models, it is possible to achieve caption quality on par with GPT-4.1 in various domains with an 89.5% reduction in costs. By leveraging CapFlow as the data synthesizer, we produce high-quality visual captions from image and video domains at scale, and obtain a generalist visual captioner via fine-tuning, namely MetaCaptioner. Through extensive experiments, we show that MetaCaptioner not only achieves comparable captioning capabilities with commercial models but also reaches top-tier multimodal performance in the open-source community. We hope CapFlow and MetaCaptioner can benefit future multimodal research by providing a strong and cost-effective visual captioning solution.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12126
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
Lei, Zhenxin
Gao, Zhangwei
Tian, Changyao
Cui, Erfei
Chen, Guanzhou
Yang, Danni
Duan, Yuchen
Wang, Zhaokai
Li, Wenhao
Wang, Weiyun
Zhao, Xiangyu
Ji, Jiayi
Qiao, Yu
Wang, Wenhai
Luo, Gen
Computer Vision and Pattern Recognition
Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains. In this task, current open-source models present a large performance gap with commercial ones, which limits various applications such as data synthesis. To bridge the gap, this paper proposes CapFlow, a novel multi-agent collaboration workflow. CapFlow demonstrates for the first time that, by capitalizing on open-source models, it is possible to achieve caption quality on par with GPT-4.1 in various domains with an 89.5% reduction in costs. By leveraging CapFlow as the data synthesizer, we produce high-quality visual captions from image and video domains at scale, and obtain a generalist visual captioner via fine-tuning, namely MetaCaptioner. Through extensive experiments, we show that MetaCaptioner not only achieves comparable captioning capabilities with commercial models but also reaches top-tier multimodal performance in the open-source community. We hope CapFlow and MetaCaptioner can benefit future multimodal research by providing a strong and cost-effective visual captioning solution.
title MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.12126