IF-VidCap: Can Video Caption Models Follow Instructions?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Shihao, Zhang, Yuanxing, Wu, Jiangtao, Lei, Zhide, He, Yiwen, Wen, Runzhe, Liao, Chenxi, Jiang, Chengkang, Ping, An, Gao, Shuo, Wang, Suhan, Bian, Zhaozhou, Zhou, Zijun, Xie, Jingyi, Zhou, Jiayi, Wang, Jing, Yao, Yifan, Xie, Weihao, Tan, Yingshui, Wang, Yanghai, Xie, Qianqian, Zhang, Zhaoxiang, Liu, Jiaheng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908604928360448
author Li, Shihao
Zhang, Yuanxing
Wu, Jiangtao
Lei, Zhide
He, Yiwen
Wen, Runzhe
Liao, Chenxi
Jiang, Chengkang
Ping, An
Gao, Shuo
Wang, Suhan
Bian, Zhaozhou
Zhou, Zijun
Xie, Jingyi
Zhou, Jiayi
Wang, Jing
Yao, Yifan
Xie, Weihao
Tan, Yingshui
Wang, Yanghai
Xie, Qianqian
Zhang, Zhaoxiang
Liu, Jiaheng
author_facet Li, Shihao
Zhang, Yuanxing
Wu, Jiangtao
Lei, Zhide
He, Yiwen
Wen, Runzhe
Liao, Chenxi
Jiang, Chengkang
Ping, An
Gao, Shuo
Wang, Suhan
Bian, Zhaozhou
Zhou, Zijun
Xie, Jingyi
Zhou, Jiayi
Wang, Jing
Yao, Yifan
Xie, Weihao
Tan, Yingshui
Wang, Yanghai
Xie, Qianqian
Zhang, Zhaoxiang
Liu, Jiaheng
contents Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained descriptions. Current benchmarks, however, primarily assess descriptive comprehensiveness while largely overlooking instruction-following capabilities. To address this gap, we introduce IF-VidCap, a new benchmark for evaluating controllable video captioning, which contains 1,400 high-quality samples. Distinct from existing video captioning or general instruction-following benchmarks, IF-VidCap incorporates a systematic framework that assesses captions on two dimensions: format correctness and content correctness. Our comprehensive evaluation of over 20 prominent models reveals a nuanced landscape: despite the continued dominance of proprietary models, the performance gap is closing, with top-tier open-source solutions now achieving near-parity. Furthermore, we find that models specialized for dense captioning underperform general-purpose MLLMs on complex instructions, indicating that future work should simultaneously advance both descriptive richness and instruction-following fidelity.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IF-VidCap: Can Video Caption Models Follow Instructions?
Li, Shihao
Zhang, Yuanxing
Wu, Jiangtao
Lei, Zhide
He, Yiwen
Wen, Runzhe
Liao, Chenxi
Jiang, Chengkang
Ping, An
Gao, Shuo
Wang, Suhan
Bian, Zhaozhou
Zhou, Zijun
Xie, Jingyi
Zhou, Jiayi
Wang, Jing
Yao, Yifan
Xie, Weihao
Tan, Yingshui
Wang, Yanghai
Xie, Qianqian
Zhang, Zhaoxiang
Liu, Jiaheng
Computer Vision and Pattern Recognition
Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained descriptions. Current benchmarks, however, primarily assess descriptive comprehensiveness while largely overlooking instruction-following capabilities. To address this gap, we introduce IF-VidCap, a new benchmark for evaluating controllable video captioning, which contains 1,400 high-quality samples. Distinct from existing video captioning or general instruction-following benchmarks, IF-VidCap incorporates a systematic framework that assesses captions on two dimensions: format correctness and content correctness. Our comprehensive evaluation of over 20 prominent models reveals a nuanced landscape: despite the continued dominance of proprietary models, the performance gap is closing, with top-tier open-source solutions now achieving near-parity. Furthermore, we find that models specialized for dense captioning underperform general-purpose MLLMs on complex instructions, indicating that future work should simultaneously advance both descriptive richness and instruction-following fidelity.
title IF-VidCap: Can Video Caption Models Follow Instructions?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.18726