Interaction Region Visual Transformer for Egocentric Action Anticipation
Fuente:
arXiv
Saved in:
| Main Authors: | Roy, Debaditya, Rajendiran, Ramanathan, Fernando, Basura |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
by: Rajendiran, Ramanathan, et al.
Published: (2023)
by: Rajendiran, Ramanathan, et al.
Published: (2023)
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
by: Rajendiran, Ramanathan, et al.
Published: (2025)
by: Rajendiran, Ramanathan, et al.
Published: (2025)
Predicting the Next Action by Modeling the Abstract Goal
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
by: Verma, Dhruv, et al.
Published: (2024)
by: Verma, Dhruv, et al.
Published: (2024)
Improving Temporal Action Segmentation via Constraint-Aware Decoding
by: Ee, Yeo Keat, et al.
Published: (2026)
by: Ee, Yeo Keat, et al.
Published: (2026)
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
by: Jaiswal, Shantanu, et al.
Published: (2024)
by: Jaiswal, Shantanu, et al.
Published: (2024)
RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
by: Zhang, Hao, et al.
Published: (2023)
by: Zhang, Hao, et al.
Published: (2023)
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
by: Leonardi, Rosario, et al.
Published: (2026)
by: Leonardi, Rosario, et al.
Published: (2026)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
by: Lai, Bolin, et al.
Published: (2023)
by: Lai, Bolin, et al.
Published: (2023)
Learning to Visually Connect Actions and their Effects
by: Parmar, Paritosh, et al.
Published: (2024)
by: Parmar, Paritosh, et al.
Published: (2024)
Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
by: Chu, Qiaohui, et al.
Published: (2025)
by: Chu, Qiaohui, et al.
Published: (2025)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
Inferring Past Human Actions in Homes with Abductive Reasoning
by: Tan, Clement, et al.
Published: (2022)
by: Tan, Clement, et al.
Published: (2022)
Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos
by: Materia, Daniele, et al.
Published: (2026)
by: Materia, Daniele, et al.
Published: (2026)
Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
by: Keat, Ee Yeo, et al.
Published: (2024)
by: Keat, Ee Yeo, et al.
Published: (2024)
Bidirectional Progressive Transformer for Interaction Intention Anticipation
by: Zhang, Zichen, et al.
Published: (2024)
by: Zhang, Zichen, et al.
Published: (2024)
GLaRE: A Graph-based Landmark Region Embedding Network for Emotion Recognition
by: Maji, Debasis, et al.
Published: (2025)
by: Maji, Debasis, et al.
Published: (2025)
Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning
by: Bangde, Yashwant Pravinrao, et al.
Published: (2026)
by: Bangde, Yashwant Pravinrao, et al.
Published: (2026)
See It Before You Grab It: Deep Learning-based Action Anticipation in Basketball
by: Roy, Arnau Barrera, et al.
Published: (2025)
by: Roy, Arnau Barrera, et al.
Published: (2025)
Action-Guided Attention for Video Action Anticipation
by: Tai, Tsung-Ming, et al.
Published: (2026)
by: Tai, Tsung-Ming, et al.
Published: (2026)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Multimodal Large Models Are Effective Action Anticipators
by: Wang, Binglu, et al.
Published: (2025)
by: Wang, Binglu, et al.
Published: (2025)
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
by: Kundu, Sanjoy, et al.
Published: (2024)
by: Kundu, Sanjoy, et al.
Published: (2024)
Interact with me: Joint Egocentric Forecasting of Intent to Interact, Attitude and Social Actions
by: Bian, Tongfei, et al.
Published: (2024)
by: Bian, Tongfei, et al.
Published: (2024)
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
A Survey on Deep Learning Techniques for Action Anticipation
by: Zhong, Zeyun, et al.
Published: (2023)
by: Zhong, Zeyun, et al.
Published: (2023)
Domain Generalization using Action Sequences for Egocentric Action Recognition
by: Nasirimajd, Amirshayan, et al.
Published: (2025)
by: Nasirimajd, Amirshayan, et al.
Published: (2025)
Situational Scene Graph for Structured Human-centric Situation Understanding
by: Sugandhika, Chinthani, et al.
Published: (2024)
by: Sugandhika, Chinthani, et al.
Published: (2024)
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Intention Action Anticipation Model with Guide-Feedback Loop Mechanism
by: Ma, Zongnan, et al.
Published: (2024)
by: Ma, Zongnan, et al.
Published: (2024)
Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
by: Sato, Yuji, et al.
Published: (2025)
by: Sato, Yuji, et al.
Published: (2025)
Human Action Anticipation: A Survey
by: Lai, Bolin, et al.
Published: (2024)
by: Lai, Bolin, et al.
Published: (2024)
Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
by: Kundu, Sanjoy, et al.
Published: (2023)
by: Kundu, Sanjoy, et al.
Published: (2023)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
Edge-Efficient Image Restoration: Transformer Distillation into State-Space Models
by: Miriyala, Srinivas Soumitri, et al.
Published: (2026)
by: Miriyala, Srinivas Soumitri, et al.
Published: (2026)
HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation
by: Wang, Zirui, et al.
Published: (2024)
by: Wang, Zirui, et al.
Published: (2024)
EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views
by: Yang, Yuhang, et al.
Published: (2024)
by: Yang, Yuhang, et al.
Published: (2024)
EgoPrompt: Prompt Learning for Egocentric Action Recognition
by: Lyu, Huaihai, et al.
Published: (2025)
by: Lyu, Huaihai, et al.
Published: (2025)
Similar Items
-
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
by: Rajendiran, Ramanathan, et al.
Published: (2023) -
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
by: Rajendiran, Ramanathan, et al.
Published: (2025) -
Predicting the Next Action by Modeling the Abstract Goal
by: Roy, Debaditya, et al.
Published: (2022) -
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
by: Verma, Dhruv, et al.
Published: (2024) -
Improving Temporal Action Segmentation via Constraint-Aware Decoding
by: Ee, Yeo Keat, et al.
Published: (2026)