An Integrated Framework for Multi-Granular Explanation of Video Summarization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tsigos, Konstantinos, Apostolidis, Evlampios, Mezaris, Vasileios
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911878688538624
author Tsigos, Konstantinos
Apostolidis, Evlampios
Mezaris, Vasileios
author_facet Tsigos, Konstantinos
Apostolidis, Evlampios
Mezaris, Vasileios
contents In this paper, we propose an integrated framework for multi-granular explanation of video summarization. This framework integrates methods for producing explanations both at the fragment level (indicating which video fragments influenced the most the decisions of the summarizer) and the more fine-grained visual object level (highlighting which visual objects were the most influential for the summarizer). To build this framework, we extend our previous work on this field, by investigating the use of a model-agnostic, perturbation-based approach for fragment-level explanation of the video summarization results, and introducing a new method that combines the results of video panoptic segmentation with an adaptation of a perturbation-based explanation approach to produce object-level explanations. The performance of the developed framework is evaluated using a state-of-the-art summarization method and two datasets for benchmarking video summarization. The findings of the conducted quantitative and qualitative evaluations demonstrate the ability of our framework to spot the most and least influential fragments and visual objects of the video for the summarizer, and to provide a comprehensive set of visual-based explanations about the output of the summarization process.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10082
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Integrated Framework for Multi-Granular Explanation of Video Summarization
Tsigos, Konstantinos
Apostolidis, Evlampios
Mezaris, Vasileios
Computer Vision and Pattern Recognition
Artificial Intelligence
In this paper, we propose an integrated framework for multi-granular explanation of video summarization. This framework integrates methods for producing explanations both at the fragment level (indicating which video fragments influenced the most the decisions of the summarizer) and the more fine-grained visual object level (highlighting which visual objects were the most influential for the summarizer). To build this framework, we extend our previous work on this field, by investigating the use of a model-agnostic, perturbation-based approach for fragment-level explanation of the video summarization results, and introducing a new method that combines the results of video panoptic segmentation with an adaptation of a perturbation-based explanation approach to produce object-level explanations. The performance of the developed framework is evaluated using a state-of-the-art summarization method and two datasets for benchmarking video summarization. The findings of the conducted quantitative and qualitative evaluations demonstrate the ability of our framework to spot the most and least influential fragments and visual objects of the video for the summarizer, and to provide a comprehensive set of visual-based explanations about the output of the summarization process.
title An Integrated Framework for Multi-Granular Explanation of Video Summarization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2405.10082