"Previously on ..." From Recaps to Story Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Aditya Kumar, Srivastava, Dhruv, Tapaswi, Makarand
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916251702394880
author Singh, Aditya Kumar
Srivastava, Dhruv
Tapaswi, Makarand
author_facet Singh, Aditya Kumar
Srivastava, Dhruv
Tapaswi, Makarand
contents We introduce multimodal story summarization by leveraging TV episode recaps - short video sequences interweaving key story moments from previous episodes to bring viewers up to speed. We propose PlotSnap, a dataset featuring two crime thriller TV shows with rich recaps and long episodes of 40 minutes. Story summarization labels are unlocked by matching recap shots to corresponding sub-stories in the episode. We propose a hierarchical model TaleSumm that processes entire episodes by creating compact shot and dialog representations, and predicts importance scores for each video shot and dialog utterance by enabling interactions between local story groups. Unlike traditional summarization, our method extracts multiple plot points from long videos. We present a thorough evaluation on story summarization, including promising cross-series generalization. TaleSumm also shows good results on classic video summarization benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_11487
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle "Previously on ..." From Recaps to Story Summarization
Singh, Aditya Kumar
Srivastava, Dhruv
Tapaswi, Makarand
Computer Vision and Pattern Recognition
We introduce multimodal story summarization by leveraging TV episode recaps - short video sequences interweaving key story moments from previous episodes to bring viewers up to speed. We propose PlotSnap, a dataset featuring two crime thriller TV shows with rich recaps and long episodes of 40 minutes. Story summarization labels are unlocked by matching recap shots to corresponding sub-stories in the episode. We propose a hierarchical model TaleSumm that processes entire episodes by creating compact shot and dialog representations, and predicts importance scores for each video shot and dialog utterance by enabling interactions between local story groups. Unlike traditional summarization, our method extracts multiple plot points from long videos. We present a thorough evaluation on story summarization, including promising cross-series generalization. TaleSumm also shows good results on classic video summarization benchmarks.
title "Previously on ..." From Recaps to Story Summarization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.11487