Multimodal Chaptering for Long-Form TV Newscast Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guetari, Khalil, Tevissen, Yannis, Petitpont, Frédéric
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916299944230912
author Guetari, Khalil
Tevissen, Yannis
Petitpont, Frédéric
author_facet Guetari, Khalil
Tevissen, Yannis
Petitpont, Frédéric
contents We propose a novel approach for automatic chaptering of TV newscast videos, addressing the challenge of structuring and organizing large collections of unsegmented broadcast content. Our method integrates both audio and visual cues through a two-stage process involving frozen neural networks and a trained LSTM network. The first stage extracts essential features from separate modalities, while the LSTM effectively fuses these features to generate accurate segment boundaries. Our proposed model has been evaluated on a diverse dataset comprising over 500 TV newscast videos of an average of 41 minutes gathered from TF1, a French TV channel, with varying lengths and topics. Experimental results demonstrate that this innovative fusion strategy achieves state of the art performance, yielding a high precision rate of 82% at IoU of 90%. Consequently, this approach significantly enhances analysis, indexing and storage capabilities for TV newscast archives, paving the way towards efficient management and utilization of vast audiovisual resources.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17590
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal Chaptering for Long-Form TV Newscast Video
Guetari, Khalil
Tevissen, Yannis
Petitpont, Frédéric
Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
We propose a novel approach for automatic chaptering of TV newscast videos, addressing the challenge of structuring and organizing large collections of unsegmented broadcast content. Our method integrates both audio and visual cues through a two-stage process involving frozen neural networks and a trained LSTM network. The first stage extracts essential features from separate modalities, while the LSTM effectively fuses these features to generate accurate segment boundaries. Our proposed model has been evaluated on a diverse dataset comprising over 500 TV newscast videos of an average of 41 minutes gathered from TF1, a French TV channel, with varying lengths and topics. Experimental results demonstrate that this innovative fusion strategy achieves state of the art performance, yielding a high precision rate of 82% at IoU of 90%. Consequently, this approach significantly enhances analysis, indexing and storage capabilities for TV newscast archives, paving the way towards efficient management and utilization of vast audiovisual resources.
title Multimodal Chaptering for Long-Form TV Newscast Video
topic Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.17590