Unsupervised Transcript-assisted Video Summarization and Highlight Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Barbakos, Spyros, Antoniadis, Charalampos, Potamianos, Gerasimos, Setti, Gianluca
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910973700341760
author Barbakos, Spyros
Antoniadis, Charalampos
Potamianos, Gerasimos
Setti, Gianluca
author_facet Barbakos, Spyros
Antoniadis, Charalampos
Potamianos, Gerasimos
Setti, Gianluca
contents Video consumption is a key part of daily life, but watching entire videos can be tedious. To address this, researchers have explored video summarization and highlight detection to identify key video segments. While some works combine video frames and transcripts, and others tackle video summarization and highlight detection using Reinforcement Learning (RL), no existing work, to the best of our knowledge, integrates both modalities within an RL framework. In this paper, we propose a multimodal pipeline that leverages video frames and their corresponding transcripts to generate a more condensed version of the video and detect highlights using a modality fusion mechanism. The pipeline is trained within an RL framework, which rewards the model for generating diverse and representative summaries while ensuring the inclusion of video segments with meaningful transcript content. The unsupervised nature of the training allows for learning from large-scale unannotated datasets, overcoming the challenge posed by the limited size of existing annotated datasets. Our experiments show that using the transcript in video summarization and highlight detection achieves superior results compared to relying solely on the visual content of the video.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23268
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Transcript-assisted Video Summarization and Highlight Detection
Barbakos, Spyros
Antoniadis, Charalampos
Potamianos, Gerasimos
Setti, Gianluca
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Video consumption is a key part of daily life, but watching entire videos can be tedious. To address this, researchers have explored video summarization and highlight detection to identify key video segments. While some works combine video frames and transcripts, and others tackle video summarization and highlight detection using Reinforcement Learning (RL), no existing work, to the best of our knowledge, integrates both modalities within an RL framework. In this paper, we propose a multimodal pipeline that leverages video frames and their corresponding transcripts to generate a more condensed version of the video and detect highlights using a modality fusion mechanism. The pipeline is trained within an RL framework, which rewards the model for generating diverse and representative summaries while ensuring the inclusion of video segments with meaningful transcript content. The unsupervised nature of the training allows for learning from large-scale unannotated datasets, overcoming the challenge posed by the limited size of existing annotated datasets. Our experiments show that using the transcript in video summarization and highlight detection achieves superior results compared to relying solely on the visual content of the video.
title Unsupervised Transcript-assisted Video Summarization and Highlight Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2505.23268