IF-VidCap: Can Video Caption Models Follow Instructions?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908604928360448 |
|---|---|
| author | Li, Shihao Zhang, Yuanxing Wu, Jiangtao Lei, Zhide He, Yiwen Wen, Runzhe Liao, Chenxi Jiang, Chengkang Ping, An Gao, Shuo Wang, Suhan Bian, Zhaozhou Zhou, Zijun Xie, Jingyi Zhou, Jiayi Wang, Jing Yao, Yifan Xie, Weihao Tan, Yingshui Wang, Yanghai Xie, Qianqian Zhang, Zhaoxiang Liu, Jiaheng |
| author_facet | Li, Shihao Zhang, Yuanxing Wu, Jiangtao Lei, Zhide He, Yiwen Wen, Runzhe Liao, Chenxi Jiang, Chengkang Ping, An Gao, Shuo Wang, Suhan Bian, Zhaozhou Zhou, Zijun Xie, Jingyi Zhou, Jiayi Wang, Jing Yao, Yifan Xie, Weihao Tan, Yingshui Wang, Yanghai Xie, Qianqian Zhang, Zhaoxiang Liu, Jiaheng |
| contents | Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained descriptions. Current benchmarks, however, primarily assess descriptive comprehensiveness while largely overlooking instruction-following capabilities. To address this gap, we introduce IF-VidCap, a new benchmark for evaluating controllable video captioning, which contains 1,400 high-quality samples. Distinct from existing video captioning or general instruction-following benchmarks, IF-VidCap incorporates a systematic framework that assesses captions on two dimensions: format correctness and content correctness. Our comprehensive evaluation of over 20 prominent models reveals a nuanced landscape: despite the continued dominance of proprietary models, the performance gap is closing, with top-tier open-source solutions now achieving near-parity. Furthermore, we find that models specialized for dense captioning underperform general-purpose MLLMs on complex instructions, indicating that future work should simultaneously advance both descriptive richness and instruction-following fidelity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_18726 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | IF-VidCap: Can Video Caption Models Follow Instructions? Li, Shihao Zhang, Yuanxing Wu, Jiangtao Lei, Zhide He, Yiwen Wen, Runzhe Liao, Chenxi Jiang, Chengkang Ping, An Gao, Shuo Wang, Suhan Bian, Zhaozhou Zhou, Zijun Xie, Jingyi Zhou, Jiayi Wang, Jing Yao, Yifan Xie, Weihao Tan, Yingshui Wang, Yanghai Xie, Qianqian Zhang, Zhaoxiang Liu, Jiaheng Computer Vision and Pattern Recognition Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained descriptions. Current benchmarks, however, primarily assess descriptive comprehensiveness while largely overlooking instruction-following capabilities. To address this gap, we introduce IF-VidCap, a new benchmark for evaluating controllable video captioning, which contains 1,400 high-quality samples. Distinct from existing video captioning or general instruction-following benchmarks, IF-VidCap incorporates a systematic framework that assesses captions on two dimensions: format correctness and content correctness. Our comprehensive evaluation of over 20 prominent models reveals a nuanced landscape: despite the continued dominance of proprietary models, the performance gap is closing, with top-tier open-source solutions now achieving near-parity. Furthermore, we find that models specialized for dense captioning underperform general-purpose MLLMs on complex instructions, indicating that future work should simultaneously advance both descriptive richness and instruction-following fidelity. |
| title | IF-VidCap: Can Video Caption Models Follow Instructions? |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.18726 |