Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Knab, Patrick, Marton, Sascha, Schubert, Philipp J., Guggiana, Drago, Bartelt, Christian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913114677575680
author Knab, Patrick
Marton, Sascha
Schubert, Philipp J.
Guggiana, Drago
Bartelt, Christian
author_facet Knab, Patrick
Marton, Sascha
Schubert, Philipp J.
Guggiana, Drago
Bartelt, Christian
contents Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but extending this paradigm to video remains challenging due to the difficulty of extracting concepts and modeling them over time. In this paper, we introduce MoTIF (Moving Temporal Interpretable Framework), a transformer-based concept architecture that operates on sequences of temporally grounded concept activations, by employing per-concept temporal self-attention to model when individual concepts recur and how their temporal patterns contribute to predictions. Central to the framework is a class-conditioned VLM-based concept discovery module that extracts object- and action-centric textual concepts from training videos, yielding temporally expressive concept sets without manual concept annotation. Across multiple video benchmarks, this combination improves over global concept bottlenecks and remains competitive within the interpretable concept-bottleneck setting, while narrowing the gap to strong black-box video baselines that we report as contextual references. Code available at github.com/patrick-knab/MoTIF.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20899
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification
Knab, Patrick
Marton, Sascha
Schubert, Philipp J.
Guggiana, Drago
Bartelt, Christian
Computer Vision and Pattern Recognition
Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but extending this paradigm to video remains challenging due to the difficulty of extracting concepts and modeling them over time. In this paper, we introduce MoTIF (Moving Temporal Interpretable Framework), a transformer-based concept architecture that operates on sequences of temporally grounded concept activations, by employing per-concept temporal self-attention to model when individual concepts recur and how their temporal patterns contribute to predictions. Central to the framework is a class-conditioned VLM-based concept discovery module that extracts object- and action-centric textual concepts from training videos, yielding temporally expressive concept sets without manual concept annotation. Across multiple video benchmarks, this combination improves over global concept bottlenecks and remains competitive within the interpretable concept-bottleneck setting, while narrowing the gap to strong black-box video baselines that we report as contextual references. Code available at github.com/patrick-knab/MoTIF.
title Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.20899