GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Wan, Zhu, Feng, Zeng, Yihan, Guo, Yuanfan, Liu, Ming, Xu, Hang, Zuo, Wangmeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909786995425280
author Xu, Wan
Zhu, Feng
Zeng, Yihan
Guo, Yuanfan
Liu, Ming
Xu, Hang
Zuo, Wangmeng
author_facet Xu, Wan
Zhu, Feng
Zeng, Yihan
Guo, Yuanfan
Liu, Ming
Xu, Hang
Zuo, Wangmeng
contents Video detailed captioning aims to generate comprehensive video descriptions to facilitate video understanding. Recently, most efforts in the video detailed captioning community have been made towards a local-to-global paradigm, which first generates local captions from video clips and then summarizes them into a global caption. However, we find this paradigm leads to less detailed and contextual-inconsistent captions, which can be attributed to (1) no mechanism to ensure fine-grained captions, and (2) weak interaction between local and global captions. To remedy the above two issues, we propose GLaVE-Cap, a Global-Local aligned framework with Vision Expert integration for Captioning, which consists of two core modules: TrackFusion enables comprehensive local caption generation, by leveraging vision experts to acquire cross-frame visual prompts, coupled with a dual-stream structure; while CaptionBridge establishes a local-global interaction, by using global context to guide local captioning, and adaptively summarizing local captions into a coherent global caption. Besides, we construct GLaVE-Bench, a comprehensive video captioning benchmark featuring 5X more queries per video than existing benchmarks, covering diverse visual dimensions to facilitate reliable evaluation. We further provide a training dataset GLaVE-1.2M containing 16K high-quality fine-grained video captions and 1.2M related question-answer pairs. Extensive experiments on four benchmarks show that our GLaVE-Cap achieves state-of-the-art performance. Besides, the ablation studies and student model analyses further validate the effectiveness of the proposed modules and the contribution of GLaVE-1.2M to the video understanding community. The source code, model weights, benchmark, and dataset will be open-sourced.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11360
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
Xu, Wan
Zhu, Feng
Zeng, Yihan
Guo, Yuanfan
Liu, Ming
Xu, Hang
Zuo, Wangmeng
Computer Vision and Pattern Recognition
Video detailed captioning aims to generate comprehensive video descriptions to facilitate video understanding. Recently, most efforts in the video detailed captioning community have been made towards a local-to-global paradigm, which first generates local captions from video clips and then summarizes them into a global caption. However, we find this paradigm leads to less detailed and contextual-inconsistent captions, which can be attributed to (1) no mechanism to ensure fine-grained captions, and (2) weak interaction between local and global captions. To remedy the above two issues, we propose GLaVE-Cap, a Global-Local aligned framework with Vision Expert integration for Captioning, which consists of two core modules: TrackFusion enables comprehensive local caption generation, by leveraging vision experts to acquire cross-frame visual prompts, coupled with a dual-stream structure; while CaptionBridge establishes a local-global interaction, by using global context to guide local captioning, and adaptively summarizing local captions into a coherent global caption. Besides, we construct GLaVE-Bench, a comprehensive video captioning benchmark featuring 5X more queries per video than existing benchmarks, covering diverse visual dimensions to facilitate reliable evaluation. We further provide a training dataset GLaVE-1.2M containing 16K high-quality fine-grained video captions and 1.2M related question-answer pairs. Extensive experiments on four benchmarks show that our GLaVE-Cap achieves state-of-the-art performance. Besides, the ablation studies and student model analyses further validate the effectiveness of the proposed modules and the contribution of GLaVE-1.2M to the video understanding community. The source code, model weights, benchmark, and dataset will be open-sourced.
title GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.11360