From Detection to Anticipation: Online Understanding of Struggles across Various Tasks and Activities
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Shijia, Wray, Michael, Mayol-Cuevas, Walterio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EvoStruggle: A Dataset Capturing the Evolution of Struggle across Activities and Skill Levels
by: Feng, Shijia, et al.
Published: (2025)
by: Feng, Shijia, et al.
Published: (2025)
Are you Struggling? Dataset and Baselines for Struggle Determination in Assembly Videos
by: Feng, Shijia, et al.
Published: (2024)
by: Feng, Shijia, et al.
Published: (2024)
Re-localization acceleration with Medoid Silhouette Clustering
by: Zhang, Hongyi, et al.
Published: (2024)
by: Zhang, Hongyi, et al.
Published: (2024)
CULTURE3D: A Large-Scale and Diverse Dataset of Cultural Landmarks and Terrains for Gaussian-Based Scene Rendering
by: Zheng, Xinyi, et al.
Published: (2025)
by: Zheng, Xinyi, et al.
Published: (2025)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
by: Zhou, Wenqi, et al.
Published: (2025)
by: Zhou, Wenqi, et al.
Published: (2025)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
by: Zheng, Xinyi, et al.
Published: (2026)
by: Zheng, Xinyi, et al.
Published: (2026)
Uncertainty-boosted Robust Video Activity Anticipation
by: Qi, Zhaobo, et al.
Published: (2024)
by: Qi, Zhaobo, et al.
Published: (2024)
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
Video, How Do Your Tokens Merge?
by: Pollard, Sam, et al.
Published: (2025)
by: Pollard, Sam, et al.
Published: (2025)
A Video Is Not Worth a Thousand Words
by: Pollard, Sam, et al.
Published: (2025)
by: Pollard, Sam, et al.
Published: (2025)
Toward Human Understanding with Controllable Synthesis
by: Cuevas-Velasquez, Hanz, et al.
Published: (2024)
by: Cuevas-Velasquez, Hanz, et al.
Published: (2024)
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
by: Leonardi, Rosario, et al.
Published: (2026)
by: Leonardi, Rosario, et al.
Published: (2026)
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2024)
by: Seminara, Luigi, et al.
Published: (2024)
EffiPerception: an Efficient Framework for Various Perception Tasks
by: Xiang, Xinhao, et al.
Published: (2024)
by: Xiang, Xinhao, et al.
Published: (2024)
From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
by: Kong, Fei, et al.
Published: (2025)
by: Kong, Fei, et al.
Published: (2025)
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
by: Flanagan, Kevin, et al.
Published: (2025)
by: Flanagan, Kevin, et al.
Published: (2025)
SHARP: Segmentation of Hands and Arms by Range using Pseudo-Depth for Enhanced Egocentric 3D Hand Pose Estimation and Action Recognition
by: Mucha, Wiktor, et al.
Published: (2024)
by: Mucha, Wiktor, et al.
Published: (2024)
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
by: Bansal, Siddhant, et al.
Published: (2024)
by: Bansal, Siddhant, et al.
Published: (2024)
Towards Egocentric 3D Hand Pose Estimation in Unseen Domains
by: Mucha, Wiktor, et al.
Published: (2026)
by: Mucha, Wiktor, et al.
Published: (2026)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
by: Zheng, Duo, et al.
Published: (2024)
by: Zheng, Duo, et al.
Published: (2024)
A Survey on Deep Learning Techniques for Action Anticipation
by: Zhong, Zeyun, et al.
Published: (2023)
by: Zhong, Zeyun, et al.
Published: (2023)
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2025)
by: Seminara, Luigi, et al.
Published: (2025)
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
by: Zhang, Wanyue, et al.
Published: (2025)
by: Zhang, Wanyue, et al.
Published: (2025)
Accident Anticipation via Temporal Occurrence Prediction
by: Zhao, Tianhao, et al.
Published: (2025)
by: Zhao, Tianhao, et al.
Published: (2025)
Multimodal Large Models Are Effective Action Anticipators
by: Wang, Binglu, et al.
Published: (2025)
by: Wang, Binglu, et al.
Published: (2025)
Bidirectional Progressive Transformer for Interaction Intention Anticipation
by: Zhang, Zichen, et al.
Published: (2024)
by: Zhang, Zichen, et al.
Published: (2024)
Anticipating Future Object Compositions without Forgetting
by: Zahran, Youssef, et al.
Published: (2024)
by: Zahran, Youssef, et al.
Published: (2024)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
Action-Guided Attention for Video Action Anticipation
by: Tai, Tsung-Ming, et al.
Published: (2026)
by: Tai, Tsung-Ming, et al.
Published: (2026)
Short-term Object Interaction Anticipation with Disentangled Object Detection @ Ego4D Short Term Object Interaction Anticipation Challenge
by: Cho, Hyunjin, et al.
Published: (2024)
by: Cho, Hyunjin, et al.
Published: (2024)
In Anticipation of Perfect Deepfake: Identity-anchored Artifact-agnostic Detection under Rebalanced Deepfake Detection Protocol
by: Wang, Wei-Han, et al.
Published: (2024)
by: Wang, Wei-Han, et al.
Published: (2024)
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
by: Chu, Qiaohui, et al.
Published: (2026)
by: Chu, Qiaohui, et al.
Published: (2026)
Real-time Traffic Accident Anticipation with Feature Reuse
by: Song, Inpyo, et al.
Published: (2025)
by: Song, Inpyo, et al.
Published: (2025)
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024)
by: Manousaki, Victoria, et al.
Published: (2024)
Interaction Region Visual Transformer for Egocentric Action Anticipation
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
PEAR: Phrase-Based Hand-Object Interaction Anticipation
by: Zhang, Zichen, et al.
Published: (2024)
by: Zhang, Zichen, et al.
Published: (2024)
VAGNet: Vision-based Accident Anticipation with Global Features
by: Vipulananthan, Vipooshan, et al.
Published: (2026)
by: Vipulananthan, Vipooshan, et al.
Published: (2026)
Similar Items
-
EvoStruggle: A Dataset Capturing the Evolution of Struggle across Activities and Skill Levels
by: Feng, Shijia, et al.
Published: (2025) -
Are you Struggling? Dataset and Baselines for Struggle Determination in Assembly Videos
by: Feng, Shijia, et al.
Published: (2024) -
Re-localization acceleration with Medoid Silhouette Clustering
by: Zhang, Hongyi, et al.
Published: (2024) -
CULTURE3D: A Large-Scale and Diverse Dataset of Cultural Landmarks and Terrains for Gaussian-Based Scene Rendering
by: Zheng, Xinyi, et al.
Published: (2025) -
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
by: Zhou, Wenqi, et al.
Published: (2025)