SUMMA: A Multimodal Large Language Model for Advertisement Summarization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Jia, Weitao, Yin, Shuo, Wen, Zhoufutu, Wang, Han, Dai, Zehui, Zhang, Kun, Li, Zhenyu, Zeng, Tao, Lv, Xiaohui
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912640119341056
author Jia, Weitao
Yin, Shuo
Wen, Zhoufutu
Wang, Han
Dai, Zehui
Zhang, Kun
Li, Zhenyu
Zeng, Tao
Lv, Xiaohui
author_facet Jia, Weitao
Yin, Shuo
Wen, Zhoufutu
Wang, Han
Dai, Zehui
Zhang, Kun
Li, Zhenyu
Zeng, Tao
Lv, Xiaohui
contents Understanding multimodal video ads is crucial for improving query-ad matching and relevance ranking on short video platforms, enhancing advertising effectiveness and user experience. However, the effective utilization of multimodal information with high commercial value still largely constrained by reliance on highly compressed video embeddings-has long been inadequate. To address this, we propose SUMMA (the abbreviation of Summarizing MultiModal Ads), a multimodal model that automatically processes video ads into summaries highlighting the content of highest commercial value, thus improving their comprehension and ranking in Douyin search-advertising systems. SUMMA is developed via a two-stage training strategy-multimodal supervised fine-tuning followed by reinforcement learning with a mixed reward mechanism-on domain-specific data containing video frames and ASR/OCR transcripts, generating commercially valuable and explainable summaries. We integrate SUMMA-generated summaries into our production pipeline, directly enhancing the candidate retrieval and relevance ranking stages in real search-advertising systems. Both offline and online experiments show substantial improvements over baselines, with online results indicating a statistically significant 1.5% increase in advertising revenue. Our work establishes a novel paradigm for condensing multimodal information into representative texts, effectively aligning visual ad content with user query intent in retrieval and recommendation scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20582
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SUMMA: A Multimodal Large Language Model for Advertisement Summarization
Jia, Weitao
Yin, Shuo
Wen, Zhoufutu
Wang, Han
Dai, Zehui
Zhang, Kun
Li, Zhenyu
Zeng, Tao
Lv, Xiaohui
Information Retrieval
Understanding multimodal video ads is crucial for improving query-ad matching and relevance ranking on short video platforms, enhancing advertising effectiveness and user experience. However, the effective utilization of multimodal information with high commercial value still largely constrained by reliance on highly compressed video embeddings-has long been inadequate. To address this, we propose SUMMA (the abbreviation of Summarizing MultiModal Ads), a multimodal model that automatically processes video ads into summaries highlighting the content of highest commercial value, thus improving their comprehension and ranking in Douyin search-advertising systems. SUMMA is developed via a two-stage training strategy-multimodal supervised fine-tuning followed by reinforcement learning with a mixed reward mechanism-on domain-specific data containing video frames and ASR/OCR transcripts, generating commercially valuable and explainable summaries. We integrate SUMMA-generated summaries into our production pipeline, directly enhancing the candidate retrieval and relevance ranking stages in real search-advertising systems. Both offline and online experiments show substantial improvements over baselines, with online results indicating a statistically significant 1.5% increase in advertising revenue. Our work establishes a novel paradigm for condensing multimodal information into representative texts, effectively aligning visual ad content with user query intent in retrieval and recommendation scenarios.
title SUMMA: A Multimodal Large Language Model for Advertisement Summarization
topic Information Retrieval
url https://arxiv.org/abs/2508.20582