Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Schoonbeek, Tim J., Hung, Shao-Hsuan, Lehman, Dan, Onvlee, Hans, Kustra, Jacek, de With, Peter H. N., van der Sommen, Fons |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Find the Assembly Mistakes: Error Segmentation for Industrial Applications
by: Lehman, Dan, et al.
Published: (2024)
by: Lehman, Dan, et al.
Published: (2024)
Supervised Representation Learning towards Generalizable Assembly State Recognition
by: Schoonbeek, Tim J., et al.
Published: (2024)
by: Schoonbeek, Tim J., et al.
Published: (2024)
MedSymmFlow: Bridging Generative Modeling and Classification in Medical Imaging through Symmetrical Flow Matching
by: Caetano, Francisco, et al.
Published: (2025)
by: Caetano, Francisco, et al.
Published: (2025)
Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative Models
by: Caetano, Francisco, et al.
Published: (2025)
by: Caetano, Francisco, et al.
Published: (2025)
Zero-Shot Image Anomaly Detection Using Generative Foundation Models
by: Abdi, Lemar, et al.
Published: (2025)
by: Abdi, Lemar, et al.
Published: (2025)
MedShift: Implicit Conditional Transport for X-Ray Domain Adaptation
by: Caetano, Francisco, et al.
Published: (2025)
by: Caetano, Francisco, et al.
Published: (2025)
Cultuur als antwoord
by: Onvlee, L.
Published: (2016)
by: Onvlee, L.
Published: (2016)
Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
by: Claessens, Cris, et al.
Published: (2025)
by: Claessens, Cris, et al.
Published: (2025)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
by: Xu, Qi'ao, et al.
Published: (2025)
by: Xu, Qi'ao, et al.
Published: (2025)
Flow-Based Generative Modeling for Optimizing Sampling Policies in Compressed Sensing Applications
by: Pavelkin, Roman, et al.
Published: (2026)
by: Pavelkin, Roman, et al.
Published: (2026)
Out-of-Distribution Detection in Medical Imaging via Diffusion Trajectories
by: Abdi, Lemar, et al.
Published: (2025)
by: Abdi, Lemar, et al.
Published: (2025)
Sub-Riemannian Snakes on the Projective Line Bundle with Applications to Segmentation of SEM Images
by: Vis, Leanne, et al.
Published: (2026)
by: Vis, Leanne, et al.
Published: (2026)
AdverX-Ray: Ensuring X-Ray Integrity Through Frequency-Sensitive Adversarial VAEs
by: Caetano, Francisco, et al.
Published: (2025)
by: Caetano, Francisco, et al.
Published: (2025)
SemiVT-Surge: Semi-Supervised Video Transformer for Surgical Phase Recognition
by: Li, Yiping, et al.
Published: (2025)
by: Li, Yiping, et al.
Published: (2025)
DisCoPatch: Taming Adversarially-driven Batch Statistics for Improved Out-of-Distribution Detection
by: Caetano, Francisco, et al.
Published: (2025)
by: Caetano, Francisco, et al.
Published: (2025)
Can Your Generative Model Detect Out-of-Distribution Covariate Shift?
by: Viviers, Christiaan, et al.
Published: (2024)
by: Viviers, Christiaan, et al.
Published: (2024)
Advancing 6-DoF Instrument Pose Estimation in Variable X-Ray Imaging Geometries
by: Viviers, Christiaan G. A., et al.
Published: (2024)
by: Viviers, Christiaan G. A., et al.
Published: (2024)
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
How to Correctly Make Mistakes: A Framework for Constructing and Benchmarking Mistake Aware Egocentric Procedural Videos
by: Loginova, Olga, et al.
Published: (2026)
by: Loginova, Olga, et al.
Published: (2026)
Recognizing of Vocal Fold Disorders From High Speed Video: Use of Spatio‐Temporal Deep Neural Networks
by: Dhouha Attia, et al.
Published: (2025)
by: Dhouha Attia, et al.
Published: (2025)
Recognizing Hand Use and Hand Role at Home After Stroke from Egocentric Video
by: Tsai, Meng-Fen, et al.
Published: (2022)
by: Tsai, Meng-Fen, et al.
Published: (2022)
SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration
by: Bohus, Dan, et al.
Published: (2025)
by: Bohus, Dan, et al.
Published: (2025)
Rocky outcrops form islands of high and unique tree biodiversity within an ocean of grass in Serengeti National Park
by: Fons van der Plas, et al.
Published: (2024)
by: Fons van der Plas, et al.
Published: (2024)
Investigating and Improving Latent Density Segmentation Models for Aleatoric Uncertainty Quantification in Medical Imaging
by: Valiuddin, M. M. Amaan, et al.
Published: (2023)
by: Valiuddin, M. M. Amaan, et al.
Published: (2023)
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2025)
by: Seminara, Luigi, et al.
Published: (2025)
Recognizing Sumsets is NP-Complete
by: Abboud, Amir, et al.
Published: (2024)
by: Abboud, Amir, et al.
Published: (2024)
Zero-Shot Temporal Interaction Localization for Egocentric Videos
by: Zhang, Erhang, et al.
Published: (2025)
by: Zhang, Erhang, et al.
Published: (2025)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
by: Xu, Zhiyang, et al.
Published: (2026)
by: Xu, Zhiyang, et al.
Published: (2026)
Deep Learning-Based Segmentation of Peritoneal Cancer Index Regions from CT Imaging
by: Gort, Pieter C., et al.
Published: (2026)
by: Gort, Pieter C., et al.
Published: (2026)
Male alternative reproductive tactics.
by: Kustra, Matthew C, et al.
Published: (2025)
by: Kustra, Matthew C, et al.
Published: (2025)
Microbes as manipulators of egg size and developmental evolution.
by: Kustra, Matthew C, et al.
Published: (2025)
by: Kustra, Matthew C, et al.
Published: (2025)
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
by: Vo, Khoa, et al.
Published: (2024)
by: Vo, Khoa, et al.
Published: (2024)
EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
by: Sun, Shitong, et al.
Published: (2026)
by: Sun, Shitong, et al.
Published: (2026)
Shaping Claims to Urban Land
by: van Overbeek, Fons
Published: (2025)
by: van Overbeek, Fons
Published: (2025)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
by: Ohkawa, Takehiko, et al.
Published: (2023)
by: Ohkawa, Takehiko, et al.
Published: (2023)
Pro$^2$Assist: Continuous Step-Aware Proactive Assistance with Multimodal Egocentric Perception for Long-Horizon Procedural Tasks
by: Xu, Lilin, et al.
Published: (2026)
by: Xu, Lilin, et al.
Published: (2026)
Context-Guided Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
by: Chetan, Aditya, et al.
Published: (2026)
by: Chetan, Aditya, et al.
Published: (2026)
Search Changes Consumers' Minds: How Recognizing Gaps Drives Sustainable Choices
by: van der Sluis, Frans, et al.
Published: (2026)
by: van der Sluis, Frans, et al.
Published: (2026)
Similar Items
-
Find the Assembly Mistakes: Error Segmentation for Industrial Applications
by: Lehman, Dan, et al.
Published: (2024) -
Supervised Representation Learning towards Generalizable Assembly State Recognition
by: Schoonbeek, Tim J., et al.
Published: (2024) -
MedSymmFlow: Bridging Generative Modeling and Classification in Medical Imaging through Symmetrical Flow Matching
by: Caetano, Francisco, et al.
Published: (2025) -
Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative Models
by: Caetano, Francisco, et al.
Published: (2025) -
Zero-Shot Image Anomaly Detection Using Generative Foundation Models
by: Abdi, Lemar, et al.
Published: (2025)