Predicting the Next Action by Modeling the Abstract Goal
Fuente:
arXiv
Saved in:
| Main Authors: | Roy, Debaditya, Fernando, Basura |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
by: Rajendiran, Ramanathan, et al.
Published: (2023)
by: Rajendiran, Ramanathan, et al.
Published: (2023)
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
by: Jaiswal, Shantanu, et al.
Published: (2024)
by: Jaiswal, Shantanu, et al.
Published: (2024)
Interaction Region Visual Transformer for Egocentric Action Anticipation
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
by: Rajendiran, Ramanathan, et al.
Published: (2025)
by: Rajendiran, Ramanathan, et al.
Published: (2025)
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
by: Verma, Dhruv, et al.
Published: (2024)
by: Verma, Dhruv, et al.
Published: (2024)
Improving Temporal Action Segmentation via Constraint-Aware Decoding
by: Ee, Yeo Keat, et al.
Published: (2026)
by: Ee, Yeo Keat, et al.
Published: (2026)
Learning to Visually Connect Actions and their Effects
by: Parmar, Paritosh, et al.
Published: (2024)
by: Parmar, Paritosh, et al.
Published: (2024)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
Generating Key Postures of Bharatanatyam Adavus with Pose Estimation
by: Kamble, Jagadish Kashinath, et al.
Published: (2026)
by: Kamble, Jagadish Kashinath, et al.
Published: (2026)
Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
by: Keat, Ee Yeo, et al.
Published: (2024)
by: Keat, Ee Yeo, et al.
Published: (2024)
DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models
by: Sufian, Abu, et al.
Published: (2025)
by: Sufian, Abu, et al.
Published: (2025)
GoalNet: Goal Areas Oriented Pedestrian Trajectory Prediction
by: Fadillah, Amar, et al.
Published: (2024)
by: Fadillah, Amar, et al.
Published: (2024)
CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes
by: Parmar, Paritosh, et al.
Published: (2024)
by: Parmar, Paritosh, et al.
Published: (2024)
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
by: Parmar, Paritosh, et al.
Published: (2025)
by: Parmar, Paritosh, et al.
Published: (2025)
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
by: Talon, Davide, et al.
Published: (2025)
by: Talon, Davide, et al.
Published: (2025)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
by: Ren, Shuhuai, et al.
Published: (2025)
by: Ren, Shuhuai, et al.
Published: (2025)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
by: Zhou, Chunting, et al.
Published: (2024)
by: Zhou, Chunting, et al.
Published: (2024)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
by: Tian, Keyu, et al.
Published: (2024)
by: Tian, Keyu, et al.
Published: (2024)
DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
by: Fekri, Pedram, et al.
Published: (2025)
by: Fekri, Pedram, et al.
Published: (2025)
Conformal Predictions for Human Action Recognition with Vision-Language Models
by: Tim, Bary, et al.
Published: (2025)
by: Tim, Bary, et al.
Published: (2025)
Grounding Video Models to Actions through Goal Conditioned Exploration
by: Luo, Yunhao, et al.
Published: (2024)
by: Luo, Yunhao, et al.
Published: (2024)
Abstract Art Interpretation Using ControlNet
by: Srivastava, Rishabh, et al.
Published: (2024)
by: Srivastava, Rishabh, et al.
Published: (2024)
Stochastic Human Motion Prediction with Memory of Action Transition and Action Characteristic
by: Tang, Jianwei, et al.
Published: (2025)
by: Tang, Jianwei, et al.
Published: (2025)
Overcoming Semantic Dilution in Transformer-Based Next Frame Prediction
by: Nguyen, Hy, et al.
Published: (2025)
by: Nguyen, Hy, et al.
Published: (2025)
Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals
by: Gillman, Nate, et al.
Published: (2026)
by: Gillman, Nate, et al.
Published: (2026)
Noise-Free Explanation for Driving Action Prediction
by: Zhu, Hongbo, et al.
Published: (2024)
by: Zhu, Hongbo, et al.
Published: (2024)
NextStop: An Improved Tracker For Panoptic LIDAR Segmentation Data
by: Alkalay, Nirit, et al.
Published: (2025)
by: Alkalay, Nirit, et al.
Published: (2025)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
Predicting Road Crossing Behaviour using Pose Detection and Sequence Modelling
by: Dasgupta, Subhasis, et al.
Published: (2025)
by: Dasgupta, Subhasis, et al.
Published: (2025)
From Attribution to Action: Jointly ALIGNing Predictions and Explanations
by: Hong, Dongsheng, et al.
Published: (2025)
by: Hong, Dongsheng, et al.
Published: (2025)
AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead
by: Chang, Aiden, et al.
Published: (2025)
by: Chang, Aiden, et al.
Published: (2025)
Fostering Video Reasoning via Next-Event Prediction
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
by: Cheng, Jen-Hao, et al.
Published: (2025)
by: Cheng, Jen-Hao, et al.
Published: (2025)
Cut2Next: Generating Next Shot via In-Context Tuning
by: He, Jingwen, et al.
Published: (2025)
by: He, Jingwen, et al.
Published: (2025)
Unsupervised Domain Adaptation for Action Recognition via Self-Ensembling and Conditional Embedding Alignment
by: Ghosh, Indrajeet, et al.
Published: (2024)
by: Ghosh, Indrajeet, et al.
Published: (2024)
SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals
by: Lin, Zihang, et al.
Published: (2026)
by: Lin, Zihang, et al.
Published: (2026)
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
by: Zhang, Huichao, et al.
Published: (2026)
by: Zhang, Huichao, et al.
Published: (2026)
MulCPred: Learning Multi-modal Concepts for Explainable Pedestrian Action Prediction
by: Feng, Yan, et al.
Published: (2024)
by: Feng, Yan, et al.
Published: (2024)
Inferring Past Human Actions in Homes with Abductive Reasoning
by: Tan, Clement, et al.
Published: (2022)
by: Tan, Clement, et al.
Published: (2022)
Similar Items
-
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
by: Rajendiran, Ramanathan, et al.
Published: (2023) -
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
by: Jaiswal, Shantanu, et al.
Published: (2024) -
Interaction Region Visual Transformer for Egocentric Action Anticipation
by: Roy, Debaditya, et al.
Published: (2022) -
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
by: Rajendiran, Ramanathan, et al.
Published: (2025) -
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
by: Verma, Dhruv, et al.
Published: (2024)