Saved in:
Bibliographic Details
Main Authors: Zhao, Fei, Pan, Da, Qi, Zelu, Shi, Ping
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.10331
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913890040807424
author Zhao, Fei
Pan, Da
Qi, Zelu
Shi, Ping
author_facet Zhao, Fei
Pan, Da
Qi, Zelu
Shi, Ping
contents In response to the rising prominence of the Metaverse, omnidirectional videos (ODVs) have garnered notable interest, gradually shifting from professional-generated content (PGC) to user-generated content (UGC). However, the study of audio-visual quality assessment (AVQA) within ODVs remains limited. To address this, we construct a dataset of UGC omnidirectional audio and video (A/V) content. The videos are captured by five individuals using two different types of omnidirectional cameras, shooting 300 videos covering 10 different scene types. A subjective AVQA experiment is conducted on the dataset to obtain the Mean Opinion Scores (MOSs) of the A/V sequences. After that, to facilitate the development of UGC-ODV AVQA fields, we construct an effective AVQA baseline model on the proposed dataset, of which the baseline model consists of video feature extraction module, audio feature extraction and audio-visual fusion module. The experimental results demonstrate that our model achieves optimal performance on the proposed dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10331
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Research on Audio-Visual Quality Assessment Dataset and Method for User-Generated Omnidirectional Video
Zhao, Fei
Pan, Da
Qi, Zelu
Shi, Ping
Computer Vision and Pattern Recognition
Image and Video Processing
In response to the rising prominence of the Metaverse, omnidirectional videos (ODVs) have garnered notable interest, gradually shifting from professional-generated content (PGC) to user-generated content (UGC). However, the study of audio-visual quality assessment (AVQA) within ODVs remains limited. To address this, we construct a dataset of UGC omnidirectional audio and video (A/V) content. The videos are captured by five individuals using two different types of omnidirectional cameras, shooting 300 videos covering 10 different scene types. A subjective AVQA experiment is conducted on the dataset to obtain the Mean Opinion Scores (MOSs) of the A/V sequences. After that, to facilitate the development of UGC-ODV AVQA fields, we construct an effective AVQA baseline model on the proposed dataset, of which the baseline model consists of video feature extraction module, audio feature extraction and audio-visual fusion module. The experimental results demonstrate that our model achieves optimal performance on the proposed dataset.
title Research on Audio-Visual Quality Assessment Dataset and Method for User-Generated Omnidirectional Video
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2506.10331