Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Xuzheng, Jiang, Chen, Zhang, Wei, Gan, Tian, Chao, Linlin, Zhao, Jianan, Cheng, Yuan, Guo, Qingpei, Chu, Wei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912380224536576
author Yu, Xuzheng
Jiang, Chen
Zhang, Wei
Gan, Tian
Chao, Linlin
Zhao, Jianan
Cheng, Yuan
Guo, Qingpei
Chu, Wei
author_facet Yu, Xuzheng
Jiang, Chen
Zhang, Wei
Gan, Tian
Chao, Linlin
Zhao, Jianan
Cheng, Yuan
Guo, Qingpei
Chu, Wei
contents With the explosive growth of video data in real-world applications, a comprehensive representation of videos becomes increasingly important. In this paper, we address the problem of video scene recognition, whose goal is to learn a high-level video representation to classify scenes in videos. Due to the diversity and complexity of video contents in realistic scenarios, this task remains a challenge. Most existing works identify scenes for videos only from visual or textual information in a temporal perspective, ignoring the valuable information hidden in single frames, while several earlier studies only recognize scenes for separate images in a non-temporal perspective. We argue that these two perspectives are both meaningful for this task and complementary to each other, meanwhile, externally introduced knowledge can also promote the comprehension of videos. We propose a novel two-stream framework to model video representations from multiple perspectives, i.e. temporal and non-temporal perspectives, and integrate the two perspectives in an end-to-end manner by self-distillation. Besides, we design a knowledge-enhanced feature fusion and label prediction method that contributes to naturally introducing knowledge into the task of video scene recognition. Experiments conducted on a real-world dataset demonstrate the effectiveness of our proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04354
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
Yu, Xuzheng
Jiang, Chen
Zhang, Wei
Gan, Tian
Chao, Linlin
Zhao, Jianan
Cheng, Yuan
Guo, Qingpei
Chu, Wei
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
With the explosive growth of video data in real-world applications, a comprehensive representation of videos becomes increasingly important. In this paper, we address the problem of video scene recognition, whose goal is to learn a high-level video representation to classify scenes in videos. Due to the diversity and complexity of video contents in realistic scenarios, this task remains a challenge. Most existing works identify scenes for videos only from visual or textual information in a temporal perspective, ignoring the valuable information hidden in single frames, while several earlier studies only recognize scenes for separate images in a non-temporal perspective. We argue that these two perspectives are both meaningful for this task and complementary to each other, meanwhile, externally introduced knowledge can also promote the comprehension of videos. We propose a novel two-stream framework to model video representations from multiple perspectives, i.e. temporal and non-temporal perspectives, and integrate the two perspectives in an end-to-end manner by self-distillation. Besides, we design a knowledge-enhanced feature fusion and label prediction method that contributes to naturally introducing knowledge into the task of video scene recognition. Experiments conducted on a real-world dataset demonstrate the effectiveness of our proposed method.
title Knowledge-enhanced Multi-perspective Video Representation Learning for Scene Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2401.04354