Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
Fuente:
arXiv
Saved in:
| Main Authors: | Benavent-Lledo, Manuel, Bacharidis, Konstantinos, Manousaki, Victoria, Papoutsakis, Konstantinos, Argyros, Antonis, Garcia-Rodriguez, Jose |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024)
by: Manousaki, Victoria, et al.
Published: (2024)
Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
by: Bacharidis, Konstantinos, et al.
Published: (2025)
by: Bacharidis, Konstantinos, et al.
Published: (2025)
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
Recognizing Unseen States of Unknown Objects by Leveraging Knowledge Graphs
by: Gouidis, Filipos, et al.
Published: (2023)
by: Gouidis, Filipos, et al.
Published: (2023)
Fusing Domain-Specific Content from Large Language Models into Knowledge Graphs for Enhanced Zero Shot Object State Classification
by: Gouidis, Filippos, et al.
Published: (2024)
by: Gouidis, Filippos, et al.
Published: (2024)
Text-driven Online Action Detection
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
ENACT: Entropy-based Clustering of Attention Input for Reducing the Computational Needs of Object Detection Transformers
by: Savathrakis, Giorgos, et al.
Published: (2024)
by: Savathrakis, Giorgos, et al.
Published: (2024)
A vision-based framework for human behavior understanding in industrial assembly lines
by: Papoutsakis, Konstantinos, et al.
Published: (2024)
by: Papoutsakis, Konstantinos, et al.
Published: (2024)
AIFloodSense: A Global Aerial Imagery Dataset for Semantic Segmentation and Understanding of Flooded Environments
by: Simantiris, Georgios, et al.
Published: (2025)
by: Simantiris, Georgios, et al.
Published: (2025)
D-PoSE: Depth as an Intermediate Representation for 3D Human Pose and Shape Estimation
by: Vasilikopoulos, Nikolaos, et al.
Published: (2024)
by: Vasilikopoulos, Nikolaos, et al.
Published: (2024)
OCCAM: Class-Agnostic, Training-Free, Prior-Free and Multi-Class Object Counting
by: Spanakis, Michail, et al.
Published: (2026)
by: Spanakis, Michail, et al.
Published: (2026)
Action-Guided Attention for Video Action Anticipation
by: Tai, Tsung-Ming, et al.
Published: (2026)
by: Tai, Tsung-Ming, et al.
Published: (2026)
Combining Facial Videos and Biosignals for Stress Estimation During Driving
by: Valergaki, Paraskevi, et al.
Published: (2026)
by: Valergaki, Paraskevi, et al.
Published: (2026)
Multimodal Large Models Are Effective Action Anticipators
by: Wang, Binglu, et al.
Published: (2025)
by: Wang, Binglu, et al.
Published: (2025)
Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images
by: Qammaz, Ammar, et al.
Published: (2024)
by: Qammaz, Ammar, et al.
Published: (2024)
AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?
by: Zhao, Qi, et al.
Published: (2023)
by: Zhao, Qi, et al.
Published: (2023)
Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos
by: Rodriguez-Juan, Javier, et al.
Published: (2025)
by: Rodriguez-Juan, Javier, et al.
Published: (2025)
Detecting Facial Image Manipulations with Multi-Layer CNN Models
by: Montejano, Alejandro Marco, et al.
Published: (2024)
by: Montejano, Alejandro Marco, et al.
Published: (2024)
Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language Models
by: Mittal, Himangi, et al.
Published: (2024)
by: Mittal, Himangi, et al.
Published: (2024)
Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
by: Karvounas, Giorgos, et al.
Published: (2025)
by: Karvounas, Giorgos, et al.
Published: (2025)
A Survey on Deep Learning Techniques for Action Anticipation
by: Zhong, Zeyun, et al.
Published: (2023)
by: Zhong, Zeyun, et al.
Published: (2023)
Multi-task Learning For Joint Action and Gesture Recognition
by: Spathis, Konstantinos, et al.
Published: (2025)
by: Spathis, Konstantinos, et al.
Published: (2025)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
by: Chatzoudis, Gerasimos, et al.
Published: (2026)
by: Chatzoudis, Gerasimos, et al.
Published: (2026)
AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
by: Pardyl, Adam, et al.
Published: (2024)
by: Pardyl, Adam, et al.
Published: (2024)
Human Action Anticipation: A Survey
by: Lai, Bolin, et al.
Published: (2024)
by: Lai, Bolin, et al.
Published: (2024)
What really matters for person re-identification? A Mixture-of-Experts Framework for Semantic Attribute Importance
by: Psalta, Athena, et al.
Published: (2025)
by: Psalta, Athena, et al.
Published: (2025)
Semantically Guided Action Anticipation
by: Diko, Anxhelo, et al.
Published: (2024)
by: Diko, Anxhelo, et al.
Published: (2024)
Action Anticipation from SoccerNet Football Video Broadcasts
by: Dalal, Mohamad, et al.
Published: (2025)
by: Dalal, Mohamad, et al.
Published: (2025)
Interaction Region Visual Transformer for Egocentric Action Anticipation
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Minimalistic Video Saliency Prediction via Efficient Decoder & Spatio Temporal Action Cues
by: Girmaji, Rohit, et al.
Published: (2025)
by: Girmaji, Rohit, et al.
Published: (2025)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
by: Wang, Yubin, et al.
Published: (2024)
by: Wang, Yubin, et al.
Published: (2024)
Intelligent Sampling Consensus for Homography Estimation in Football Videos Using Featureless Unpaired Points
by: Nousias, George, et al.
Published: (2023)
by: Nousias, George, et al.
Published: (2023)
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
by: Wu, Zhixuan, et al.
Published: (2026)
by: Wu, Zhixuan, et al.
Published: (2026)
Multi-level and Multi-modal Action Anticipation
by: Kim, Seulgi, et al.
Published: (2025)
by: Kim, Seulgi, et al.
Published: (2025)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
Uncertainty-boosted Robust Video Activity Anticipation
by: Qi, Zhaobo, et al.
Published: (2024)
by: Qi, Zhaobo, et al.
Published: (2024)
From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Intention Action Anticipation Model with Guide-Feedback Loop Mechanism
by: Ma, Zongnan, et al.
Published: (2024)
by: Ma, Zongnan, et al.
Published: (2024)
Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
by: Sato, Yuji, et al.
Published: (2025)
by: Sato, Yuji, et al.
Published: (2025)
Similar Items
-
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
by: Benavent-Lledo, Manuel, et al.
Published: (2026) -
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024) -
Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
by: Bacharidis, Konstantinos, et al.
Published: (2025) -
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
by: Benavent-Lledo, Manuel, et al.
Published: (2024) -
Recognizing Unseen States of Unknown Objects by Leveraging Knowledge Graphs
by: Gouidis, Filipos, et al.
Published: (2023)