Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
Fuente:
arXiv
Saved in:
| Main Authors: | Bacharidis, Konstantinos, Argyros, Antonis A. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024)
by: Manousaki, Victoria, et al.
Published: (2024)
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
ENACT: Entropy-based Clustering of Attention Input for Reducing the Computational Needs of Object Detection Transformers
by: Savathrakis, Giorgos, et al.
Published: (2024)
by: Savathrakis, Giorgos, et al.
Published: (2024)
Recognizing Unseen States of Unknown Objects by Leveraging Knowledge Graphs
by: Gouidis, Filipos, et al.
Published: (2023)
by: Gouidis, Filipos, et al.
Published: (2023)
D-PoSE: Depth as an Intermediate Representation for 3D Human Pose and Shape Estimation
by: Vasilikopoulos, Nikolaos, et al.
Published: (2024)
by: Vasilikopoulos, Nikolaos, et al.
Published: (2024)
OCCAM: Class-Agnostic, Training-Free, Prior-Free and Multi-Class Object Counting
by: Spanakis, Michail, et al.
Published: (2026)
by: Spanakis, Michail, et al.
Published: (2026)
Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images
by: Qammaz, Ammar, et al.
Published: (2024)
by: Qammaz, Ammar, et al.
Published: (2024)
AIFloodSense: A Global Aerial Imagery Dataset for Semantic Segmentation and Understanding of Flooded Environments
by: Simantiris, Georgios, et al.
Published: (2025)
by: Simantiris, Georgios, et al.
Published: (2025)
How to Correctly Make Mistakes: A Framework for Constructing and Benchmarking Mistake Aware Egocentric Procedural Videos
by: Loginova, Olga, et al.
Published: (2026)
by: Loginova, Olga, et al.
Published: (2026)
Procedural Mistake Detection via Action Effect Modeling
by: Guo, Wenliang, et al.
Published: (2025)
by: Guo, Wenliang, et al.
Published: (2025)
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2024)
by: Seminara, Luigi, et al.
Published: (2024)
Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
by: Karvounas, Giorgos, et al.
Published: (2025)
by: Karvounas, Giorgos, et al.
Published: (2025)
Combining Facial Videos and Biosignals for Stress Estimation During Driving
by: Valergaki, Paraskevi, et al.
Published: (2026)
by: Valergaki, Paraskevi, et al.
Published: (2026)
MistExit: Learning to Exit for Early Mistake Detection in Procedural Videos
by: Majumder, Sagnik, et al.
Published: (2026)
by: Majumder, Sagnik, et al.
Published: (2026)
Fusing Domain-Specific Content from Large Language Models into Knowledge Graphs for Enhanced Zero Shot Object State Classification
by: Gouidis, Filippos, et al.
Published: (2024)
by: Gouidis, Filippos, et al.
Published: (2024)
Technical Report for Egocentric Mistake Detection for the HoloAssist Challenge
by: Patsch, Constantin, et al.
Published: (2025)
by: Patsch, Constantin, et al.
Published: (2025)
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
by: Mazzamuto, Michele, et al.
Published: (2024)
by: Mazzamuto, Michele, et al.
Published: (2024)
EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
by: Haneji, Yuto, et al.
Published: (2024)
by: Haneji, Yuto, et al.
Published: (2024)
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
by: Zhang, Zory, et al.
Published: (2025)
by: Zhang, Zory, et al.
Published: (2025)
Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Find the Assembly Mistakes: Error Segmentation for Industrial Applications
by: Lehman, Dan, et al.
Published: (2024)
by: Lehman, Dan, et al.
Published: (2024)
Making Better Mistakes in CLIP-Based Zero-Shot Classification with Hierarchy-Aware Language Prompts
by: Liang, Tong, et al.
Published: (2025)
by: Liang, Tong, et al.
Published: (2025)
A Review of 3D Object Detection with Vision-Language Models
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Vision Foundation Models in Medical Image Analysis: Advances and Challenges
by: Liang, Pengchen, et al.
Published: (2025)
by: Liang, Pengchen, et al.
Published: (2025)
Deep Learning Advances in Vision-Based Traffic Accident Anticipation: A Comprehensive Review of Methods, Datasets, and Future Directions
by: Lin, Ruonan, et al.
Published: (2025)
by: Lin, Ruonan, et al.
Published: (2025)
CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities
by: Peddi, Rohith, et al.
Published: (2023)
by: Peddi, Rohith, et al.
Published: (2023)
Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
by: Li, Junlong, et al.
Published: (2025)
by: Li, Junlong, et al.
Published: (2025)
Challenges, Advances, and Evaluation Metrics in Medical Image Enhancement: A Systematic Literature Review
by: Chin, Chun Wai, et al.
Published: (2025)
by: Chin, Chun Wai, et al.
Published: (2025)
Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training
by: Chen, Xinyan, et al.
Published: (2023)
by: Chen, Xinyan, et al.
Published: (2023)
CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning
by: Shi, Lei, et al.
Published: (2025)
by: Shi, Lei, et al.
Published: (2025)
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2025)
by: Seminara, Luigi, et al.
Published: (2025)
Vision-Based Anti Unmanned Aerial Technology: Opportunities and Challenges
by: Ding, Guanghai, et al.
Published: (2025)
by: Ding, Guanghai, et al.
Published: (2025)
Visual Hand Gesture Recognition with Deep Learning: A Comprehensive Review of Methods, Datasets, Challenges and Future Research Directions
by: Foteinos, Konstantinos, et al.
Published: (2025)
by: Foteinos, Konstantinos, et al.
Published: (2025)
Class-wise Autoencoders Measure Classification Difficulty And Detect Label Mistakes
by: Marks, Jacob, et al.
Published: (2024)
by: Marks, Jacob, et al.
Published: (2024)
TI-PREGO: Chain of Thought and In-Context Learning for Online Mistake Detection in PRocedural EGOcentric Videos
by: Plini, Leonardo, et al.
Published: (2024)
by: Plini, Leonardo, et al.
Published: (2024)
Distilling Vision Transformers for Distortion-Robust Representation Learning
by: Alexis, Konstantinos, et al.
Published: (2026)
by: Alexis, Konstantinos, et al.
Published: (2026)
A Vision Language Model for Generating Procedural Plant Architecture Representations from Simulated Images
by: Yun, Heesup, et al.
Published: (2026)
by: Yun, Heesup, et al.
Published: (2026)
Learning from Mistakes: Loss-Aware Memory Enhanced Continual Learning for LiDAR Place Recognition
by: Wang, Xufei, et al.
Published: (2025)
by: Wang, Xufei, et al.
Published: (2025)
Similar Items
-
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024) -
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
by: Benavent-Lledo, Manuel, et al.
Published: (2026) -
Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
by: Benavent-Lledo, Manuel, et al.
Published: (2025) -
ENACT: Entropy-based Clustering of Attention Input for Reducing the Computational Needs of Object Detection Transformers
by: Savathrakis, Giorgos, et al.
Published: (2024) -
Recognizing Unseen States of Unknown Objects by Leveraging Knowledge Graphs
by: Gouidis, Filipos, et al.
Published: (2023)