When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gonzálbez-Biosca, Daniel, Cabacas-Maso, Josep, Ventura, Carles, Benito-Altamirano, Ismael
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908579244539904
author Gonzálbez-Biosca, Daniel
Cabacas-Maso, Josep
Ventura, Carles
Benito-Altamirano, Ismael
author_facet Gonzálbez-Biosca, Daniel
Cabacas-Maso, Josep
Ventura, Carles
Benito-Altamirano, Ismael
contents Automated video editing remains an underexplored task in the computer vision and multimedia domains, especially when contrasted with the growing interest in video generation and scene understanding. In this work, we address the specific challenge of editing multicamera recordings of classical music concerts by decomposing the problem into two key sub-tasks: when to cut and how to cut. Building on recent literature, we propose a novel multimodal architecture for the temporal segmentation task (when to cut), which integrates log-mel spectrograms from the audio signals, plus an optional image embedding, and scalar temporal features through a lightweight convolutional-transformer pipeline. For the spatial selection task (how to cut), we improve the literature by updating from old backbones, e.g. ResNet, with a CLIP-based encoder and constraining distractor selection to segments from the same concert. Our dataset was constructed following a pseudo-labeling approach, in which raw video data was automatically clustered into coherent shot segments. We show that our models outperformed previous baselines in detecting cut points and provide competitive visual shot selection, advancing the state of the art in multimodal automated video editing.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach
Gonzálbez-Biosca, Daniel
Cabacas-Maso, Josep
Ventura, Carles
Benito-Altamirano, Ismael
Computer Vision and Pattern Recognition
Multimedia
Automated video editing remains an underexplored task in the computer vision and multimedia domains, especially when contrasted with the growing interest in video generation and scene understanding. In this work, we address the specific challenge of editing multicamera recordings of classical music concerts by decomposing the problem into two key sub-tasks: when to cut and how to cut. Building on recent literature, we propose a novel multimodal architecture for the temporal segmentation task (when to cut), which integrates log-mel spectrograms from the audio signals, plus an optional image embedding, and scalar temporal features through a lightweight convolutional-transformer pipeline. For the spatial selection task (how to cut), we improve the literature by updating from old backbones, e.g. ResNet, with a CLIP-based encoder and constraining distractor selection to segments from the same concert. Our dataset was constructed following a pseudo-labeling approach, in which raw video data was automatically clustered into coherent shot segments. We show that our models outperformed previous baselines in detecting cut points and provide competitive visual shot selection, advancing the state of the art in multimodal automated video editing.
title When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2510.05661