TI-PREGO: Chain of Thought and In-Context Learning for Online Mistake Detection in PRocedural EGOcentric Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Plini, Leonardo, Scofano, Luca, De Matteis, Edoardo, di Melendugno, Guido Maria D'Amely, Flaborea, Alessandro, Sanchietti, Andrea, Farinella, Giovanni Maria, Galasso, Fabio, Furnari, Antonino
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915708210774016
author Plini, Leonardo
Scofano, Luca
De Matteis, Edoardo
di Melendugno, Guido Maria D'Amely
Flaborea, Alessandro
Sanchietti, Andrea
Farinella, Giovanni Maria
Galasso, Fabio
Furnari, Antonino
author_facet Plini, Leonardo
Scofano, Luca
De Matteis, Edoardo
di Melendugno, Guido Maria D'Amely
Flaborea, Alessandro
Sanchietti, Andrea
Farinella, Giovanni Maria
Galasso, Fabio
Furnari, Antonino
contents Identifying procedural errors online from egocentric videos is a critical yet challenging task across various domains, including manufacturing, healthcare, and skill-based training. The nature of such mistakes is inherently open-set, as unforeseen or novel errors may occur, necessitating robust detection systems that do not rely on prior examples of failure. Currently, however, no technique effectively detects open-set procedural mistakes online. We propose a dual branch architecture to address this problem in an online fashion: one branch continuously performs step recognition from the input egocentric video, while the other anticipates future steps based on the recognition module's output. Mistakes are detected as mismatches between the currently recognized action and the action predicted by the anticipation module. The recognition branch takes input frames, predicts the current action, and aggregates frame-level results into action tokens. The anticipation branch, specifically, leverages the solid pattern-matching capabilities of Large Language Models (LLMs) to predict action tokens based on previously predicted ones. Given the online nature of the task, we also thoroughly benchmark the difficulties associated with per-frame evaluations, particularly the need for accurate and timely predictions in dynamic online scenarios. Extensive experiments on two procedural datasets demonstrate the challenges and opportunities of leveraging a dual-branch architecture for mistake detection, showcasing the effectiveness of our proposed approach. In a thorough evaluation including recognition and anticipation variants and state-of-the-art models, our method reveals its robustness and effectiveness in online applications.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02570
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TI-PREGO: Chain of Thought and In-Context Learning for Online Mistake Detection in PRocedural EGOcentric Videos
Plini, Leonardo
Scofano, Luca
De Matteis, Edoardo
di Melendugno, Guido Maria D'Amely
Flaborea, Alessandro
Sanchietti, Andrea
Farinella, Giovanni Maria
Galasso, Fabio
Furnari, Antonino
Computer Vision and Pattern Recognition
Identifying procedural errors online from egocentric videos is a critical yet challenging task across various domains, including manufacturing, healthcare, and skill-based training. The nature of such mistakes is inherently open-set, as unforeseen or novel errors may occur, necessitating robust detection systems that do not rely on prior examples of failure. Currently, however, no technique effectively detects open-set procedural mistakes online. We propose a dual branch architecture to address this problem in an online fashion: one branch continuously performs step recognition from the input egocentric video, while the other anticipates future steps based on the recognition module's output. Mistakes are detected as mismatches between the currently recognized action and the action predicted by the anticipation module. The recognition branch takes input frames, predicts the current action, and aggregates frame-level results into action tokens. The anticipation branch, specifically, leverages the solid pattern-matching capabilities of Large Language Models (LLMs) to predict action tokens based on previously predicted ones. Given the online nature of the task, we also thoroughly benchmark the difficulties associated with per-frame evaluations, particularly the need for accurate and timely predictions in dynamic online scenarios. Extensive experiments on two procedural datasets demonstrate the challenges and opportunities of leveraging a dual-branch architecture for mistake detection, showcasing the effectiveness of our proposed approach. In a thorough evaluation including recognition and anticipation variants and state-of-the-art models, our method reveals its robustness and effectiveness in online applications.
title TI-PREGO: Chain of Thought and In-Context Learning for Online Mistake Detection in PRocedural EGOcentric Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.02570