Edit As You Wish: Video Caption Editing with Multi-grained User Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Linli, Zhang, Yuanmeng, Wang, Ziheng, Hou, Xinglin, Ge, Tiezheng, Jiang, Yuning, Sun, Xu, Jin, Qin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916350009540608
author Yao, Linli
Zhang, Yuanmeng
Wang, Ziheng
Hou, Xinglin
Ge, Tiezheng
Jiang, Yuning
Sun, Xu
Jin, Qin
author_facet Yao, Linli
Zhang, Yuanmeng
Wang, Ziheng
Hou, Xinglin
Ge, Tiezheng
Jiang, Yuning
Sun, Xu
Jin, Qin
contents Automatically narrating videos in natural language complying with user requests, i.e. Controllable Video Captioning task, can help people manage massive videos with desired intentions. However, existing works suffer from two shortcomings: 1) the control signal is single-grained which can not satisfy diverse user intentions; 2) the video description is generated in a single round which can not be further edited to meet dynamic needs. In this paper, we propose a novel \textbf{V}ideo \textbf{C}aption \textbf{E}diting \textbf{(VCE)} task to automatically revise an existing video description guided by multi-grained user requests. Inspired by human writing-revision habits, we design the user command as a pivotal triplet \{\textit{operation, position, attribute}\} to cover diverse user needs from coarse-grained to fine-grained. To facilitate the VCE task, we \textit{automatically} construct an open-domain benchmark dataset named VATEX-EDIT and \textit{manually} collect an e-commerce dataset called EMMAD-EDIT. We further propose a specialized small-scale model (i.e., OPA) compared with two generalist Large Multi-modal Models to perform an exhaustive analysis of the novel task. For evaluation, we adopt comprehensive metrics considering caption fluency, command-caption consistency, and video-caption alignment. Experiments reveal the task challenges of fine-grained multi-modal semantics understanding and processing. Our datasets, codes, and evaluation tools are available at https://github.com/yaolinli/VCE.
format Preprint
id arxiv_https___arxiv_org_abs_2305_08389
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Edit As You Wish: Video Caption Editing with Multi-grained User Control
Yao, Linli
Zhang, Yuanmeng
Wang, Ziheng
Hou, Xinglin
Ge, Tiezheng
Jiang, Yuning
Sun, Xu
Jin, Qin
Computer Vision and Pattern Recognition
Multimedia
Automatically narrating videos in natural language complying with user requests, i.e. Controllable Video Captioning task, can help people manage massive videos with desired intentions. However, existing works suffer from two shortcomings: 1) the control signal is single-grained which can not satisfy diverse user intentions; 2) the video description is generated in a single round which can not be further edited to meet dynamic needs. In this paper, we propose a novel \textbf{V}ideo \textbf{C}aption \textbf{E}diting \textbf{(VCE)} task to automatically revise an existing video description guided by multi-grained user requests. Inspired by human writing-revision habits, we design the user command as a pivotal triplet \{\textit{operation, position, attribute}\} to cover diverse user needs from coarse-grained to fine-grained. To facilitate the VCE task, we \textit{automatically} construct an open-domain benchmark dataset named VATEX-EDIT and \textit{manually} collect an e-commerce dataset called EMMAD-EDIT. We further propose a specialized small-scale model (i.e., OPA) compared with two generalist Large Multi-modal Models to perform an exhaustive analysis of the novel task. For evaluation, we adopt comprehensive metrics considering caption fluency, command-caption consistency, and video-caption alignment. Experiments reveal the task challenges of fine-grained multi-modal semantics understanding and processing. Our datasets, codes, and evaluation tools are available at https://github.com/yaolinli/VCE.
title Edit As You Wish: Video Caption Editing with Multi-grained User Control
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2305.08389