AdsQA: Towards Advertisement Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Long, Xinwei, Tian, Kai, Xu, Peng, Jia, Guoli, Li, Jingxuan, Yang, Sa, Shao, Yihua, Zhang, Kaiyan, Jiang, Che, Xu, Hao, Liu, Yang, Ma, Jiaheng, Zhou, Bowen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909780320190464
author Long, Xinwei
Tian, Kai
Xu, Peng
Jia, Guoli
Li, Jingxuan
Yang, Sa
Shao, Yihua
Zhang, Kaiyan
Jiang, Che
Xu, Hao
Liu, Yang
Ma, Jiaheng
Zhou, Bowen
author_facet Long, Xinwei
Tian, Kai
Xu, Peng
Jia, Guoli
Li, Jingxuan
Yang, Sa
Shao, Yihua
Zhang, Kaiyan
Jiang, Che
Xu, Hao
Liu, Yang
Ma, Jiaheng
Zhou, Bowen
contents Large language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpose models to continuously evolve via learning deeper expertise. Now is thus the time further to extend the diversity of specialized applications for knowledgeable LLMs, though collecting high quality data with unexpected and informative tasks is challenging. In this paper, we propose to use advertisement (ad) videos as a challenging test-bed to probe the ability of LLMs in perceiving beyond the objective physical content of common visual domain. Our motivation is to take full advantage of the clue-rich and information-dense ad videos' traits, e.g., marketing logic, persuasive strategies, and audience engagement. Our contribution is three-fold: (1) To our knowledge, this is the first attempt to use ad videos with well-designed tasks to evaluate LLMs. We contribute AdsQA, a challenging ad Video QA benchmark derived from 1,544 ad videos with 10,962 clips, totaling 22.7 hours, providing 5 challenging tasks. (2) We propose ReAd-R, a Deepseek-R1 styled RL model that reflects on questions, and generates answers via reward-driven optimization. (3) We benchmark 14 top-tier LLMs on AdsQA, and our \texttt{ReAd-R}~achieves the state-of-the-art outperforming strong competitors equipped with long-chain reasoning capabilities by a clear margin.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdsQA: Towards Advertisement Video Understanding
Long, Xinwei
Tian, Kai
Xu, Peng
Jia, Guoli
Li, Jingxuan
Yang, Sa
Shao, Yihua
Zhang, Kaiyan
Jiang, Che
Xu, Hao
Liu, Yang
Ma, Jiaheng
Zhou, Bowen
Computer Vision and Pattern Recognition
Large language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpose models to continuously evolve via learning deeper expertise. Now is thus the time further to extend the diversity of specialized applications for knowledgeable LLMs, though collecting high quality data with unexpected and informative tasks is challenging. In this paper, we propose to use advertisement (ad) videos as a challenging test-bed to probe the ability of LLMs in perceiving beyond the objective physical content of common visual domain. Our motivation is to take full advantage of the clue-rich and information-dense ad videos' traits, e.g., marketing logic, persuasive strategies, and audience engagement. Our contribution is three-fold: (1) To our knowledge, this is the first attempt to use ad videos with well-designed tasks to evaluate LLMs. We contribute AdsQA, a challenging ad Video QA benchmark derived from 1,544 ad videos with 10,962 clips, totaling 22.7 hours, providing 5 challenging tasks. (2) We propose ReAd-R, a Deepseek-R1 styled RL model that reflects on questions, and generates answers via reward-driven optimization. (3) We benchmark 14 top-tier LLMs on AdsQA, and our \texttt{ReAd-R}~achieves the state-of-the-art outperforming strong competitors equipped with long-chain reasoning capabilities by a clear margin.
title AdsQA: Towards Advertisement Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.08621