VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yunhao, Wu, Sijing, Gao, Zhilin, Zhang, Zicheng, Jia, Qi, Duan, Huiyu, Min, Xiongkuo, Zhai, Guangtao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917238262464512
author Li, Yunhao
Wu, Sijing
Gao, Zhilin
Zhang, Zicheng
Jia, Qi
Duan, Huiyu
Min, Xiongkuo
Zhai, Guangtao
author_facet Li, Yunhao
Wu, Sijing
Gao, Zhilin
Zhang, Zicheng
Jia, Qi
Duan, Huiyu
Min, Xiongkuo
Zhai, Guangtao
contents Large multimodal models (LMMs) have demonstrated outstanding capabilities in various visual perception tasks, which has in turn made the evaluation of LMMs significant. However, the capability of video aesthetic quality assessment, which is a fundamental ability for human, remains underexplored for LMMs. To address this, we introduce VideoAesBench, a comprehensive benchmark for evaluating LMMs' understanding of video aesthetic quality. VideoAesBench has several significant characteristics: (1) Diverse content including 1,804 videos from multiple video sources including user-generated (UGC), AI-generated (AIGC), compressed, robotic-generated (RGC), and game videos. (2) Multiple question formats containing traditional single-choice questions, multi-choice questions, True or False questions, and a novel open-ended questions for video aesthetics description. (3) Holistic video aesthetics dimensions including visual form related questions from 5 aspects, visual style related questions from 4 aspects, and visual affectiveness questions from 3 aspects. Based on VideoAesBench, we benchmark 23 open-source and commercial large multimodal models. Our findings show that current LMMs only contain basic video aesthetics perception ability, their performance remains incomplete and imprecise. We hope our VideoAesBench can be served as a strong testbed and offer insights for explainable video aesthetics assessment. The data will be released on https://github.com/michaelliyunhao/VideoAesBench
format Preprint
id arxiv_https___arxiv_org_abs_2601_21915
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models
Li, Yunhao
Wu, Sijing
Gao, Zhilin
Zhang, Zicheng
Jia, Qi
Duan, Huiyu
Min, Xiongkuo
Zhai, Guangtao
Computer Vision and Pattern Recognition
Large multimodal models (LMMs) have demonstrated outstanding capabilities in various visual perception tasks, which has in turn made the evaluation of LMMs significant. However, the capability of video aesthetic quality assessment, which is a fundamental ability for human, remains underexplored for LMMs. To address this, we introduce VideoAesBench, a comprehensive benchmark for evaluating LMMs' understanding of video aesthetic quality. VideoAesBench has several significant characteristics: (1) Diverse content including 1,804 videos from multiple video sources including user-generated (UGC), AI-generated (AIGC), compressed, robotic-generated (RGC), and game videos. (2) Multiple question formats containing traditional single-choice questions, multi-choice questions, True or False questions, and a novel open-ended questions for video aesthetics description. (3) Holistic video aesthetics dimensions including visual form related questions from 5 aspects, visual style related questions from 4 aspects, and visual affectiveness questions from 3 aspects. Based on VideoAesBench, we benchmark 23 open-source and commercial large multimodal models. Our findings show that current LMMs only contain basic video aesthetics perception ability, their performance remains incomplete and imprecise. We hope our VideoAesBench can be served as a strong testbed and offer insights for explainable video aesthetics assessment. The data will be released on https://github.com/michaelliyunhao/VideoAesBench
title VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.21915