USV: Towards Understanding the User-generated Short-form Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Haoyue, Xu, Su, Jin, Liwei, Wu, Wayne, Qian, Chen, Wang, Limin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911700849000448
author Cheng, Haoyue
Xu, Su
Jin, Liwei
Wu, Wayne
Qian, Chen
Wang, Limin
author_facet Cheng, Haoyue
Xu, Su
Jin, Liwei
Wu, Wayne
Qian, Chen
Wang, Limin
contents Several large-scale video datasets have been published these years and have advanced the area of video understanding. However, the newly emerged user-generated short-form videos have rarely been studied. This paper presents USV, the User-generated Short-form Video dataset for high-level semantic video understanding. The dataset contains around 224K videos collected from UGC platforms by label queries without extra manual verification and trimming. Although video understanding has achieved plausible improvement these years, most works focus on instance-level recognition, which is not sufficient for learning the representation of the high-level semantic information of videos. Therefore, we further establish two tasks: topic recognition and video-text retrieval on USV. We propose two unified and effective baseline methods Multi-Modality Fusion Network (MMF-Net) and Video-Text Contrastive Learning (VTCL), to tackle the topic recognition task and video-text retrieval respectively, and carry out comprehensive benchmarks to facilitate future research. Our project page is https://usvdataset.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20838
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle USV: Towards Understanding the User-generated Short-form Videos
Cheng, Haoyue
Xu, Su
Jin, Liwei
Wu, Wayne
Qian, Chen
Wang, Limin
Computer Vision and Pattern Recognition
Artificial Intelligence
Several large-scale video datasets have been published these years and have advanced the area of video understanding. However, the newly emerged user-generated short-form videos have rarely been studied. This paper presents USV, the User-generated Short-form Video dataset for high-level semantic video understanding. The dataset contains around 224K videos collected from UGC platforms by label queries without extra manual verification and trimming. Although video understanding has achieved plausible improvement these years, most works focus on instance-level recognition, which is not sufficient for learning the representation of the high-level semantic information of videos. Therefore, we further establish two tasks: topic recognition and video-text retrieval on USV. We propose two unified and effective baseline methods Multi-Modality Fusion Network (MMF-Net) and Video-Text Contrastive Learning (VTCL), to tackle the topic recognition task and video-text retrieval respectively, and carry out comprehensive benchmarks to facilitate future research. Our project page is https://usvdataset.github.io.
title USV: Towards Understanding the User-generated Short-form Videos
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.20838