Training-free Online Video Step Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zanella, Luca, Mancini, Massimiliano, Wang, Yiming, Tonioni, Alessio, Ricci, Elisa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918163941163008
author Zanella, Luca
Mancini, Massimiliano
Wang, Yiming
Tonioni, Alessio
Ricci, Elisa
author_facet Zanella, Luca
Mancini, Massimiliano
Wang, Yiming
Tonioni, Alessio
Ricci, Elisa
contents Given a task and a set of steps composing it, Video Step Grounding (VSG) aims to detect which steps are performed in a video. Standard approaches for this task require a labeled training set (e.g., with step-level annotations or narrations), which may be costly to collect. Moreover, they process the full video offline, limiting their applications for scenarios requiring online decisions. Thus, in this work, we explore how to perform VSG online and without training. We achieve this by exploiting the zero-shot capabilities of recent Large Multimodal Models (LMMs). In particular, we use LMMs to predict the step associated with a restricted set of frames, without access to the whole video. We show that this online strategy without task-specific tuning outperforms offline and training-based models. Motivated by this finding, we develop Bayesian Grounding with Large Multimodal Models (BaGLM), further injecting knowledge of past frames into the LMM-based predictions. BaGLM exploits Bayesian filtering principles, modeling step transitions via (i) a dependency matrix extracted through large language models and (ii) an estimation of step progress. Experiments on three datasets show superior performance of BaGLM over state-of-the-art training-based offline methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16989
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training-free Online Video Step Grounding
Zanella, Luca
Mancini, Massimiliano
Wang, Yiming
Tonioni, Alessio
Ricci, Elisa
Computer Vision and Pattern Recognition
Given a task and a set of steps composing it, Video Step Grounding (VSG) aims to detect which steps are performed in a video. Standard approaches for this task require a labeled training set (e.g., with step-level annotations or narrations), which may be costly to collect. Moreover, they process the full video offline, limiting their applications for scenarios requiring online decisions. Thus, in this work, we explore how to perform VSG online and without training. We achieve this by exploiting the zero-shot capabilities of recent Large Multimodal Models (LMMs). In particular, we use LMMs to predict the step associated with a restricted set of frames, without access to the whole video. We show that this online strategy without task-specific tuning outperforms offline and training-based models. Motivated by this finding, we develop Bayesian Grounding with Large Multimodal Models (BaGLM), further injecting knowledge of past frames into the LMM-based predictions. BaGLM exploits Bayesian filtering principles, modeling step transitions via (i) a dependency matrix extracted through large language models and (ii) an estimation of step progress. Experiments on three datasets show superior performance of BaGLM over state-of-the-art training-based offline methods.
title Training-free Online Video Step Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.16989