Zero-Shot Temporal Interaction Localization for Egocentric Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Erhang, Ma, Junyi, Zheng, Yin-Dong, Zhou, Yixuan, Wang, Hesheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos
by: Ma, Junyi, et al.
Published: (2025)
by: Ma, Junyi, et al.
Published: (2025)
Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
by: Ma, Junyi, et al.
Published: (2025)
by: Ma, Junyi, et al.
Published: (2025)
Robot Learning from Human Videos: A Survey
by: Ma, Junyi, et al.
Published: (2026)
by: Ma, Junyi, et al.
Published: (2026)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
by: Wang, Qineng, et al.
Published: (2025)
by: Wang, Qineng, et al.
Published: (2025)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
by: Xiao, Junbin, et al.
Published: (2026)
by: Xiao, Junbin, et al.
Published: (2026)
Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos
by: Chen, Mingfei, et al.
Published: (2025)
by: Chen, Mingfei, et al.
Published: (2025)
SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds
by: Zhou, Yunsong, et al.
Published: (2026)
by: Zhou, Yunsong, et al.
Published: (2026)
NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
To Move or Not to Move: Constraint-based Planning Enables Zero-Shot Generalization for Interactive Navigation
by: Vashisth, Apoorva, et al.
Published: (2026)
by: Vashisth, Apoorva, et al.
Published: (2026)
HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos
by: Banerjee, Prithviraj, et al.
Published: (2024)
by: Banerjee, Prithviraj, et al.
Published: (2024)
NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning
by: Fu, Jiahui, et al.
Published: (2026)
by: Fu, Jiahui, et al.
Published: (2026)
Leveraging Unknown Objects to Construct Labeled-Unlabeled Meta-Relationships for Zero-Shot Object Navigation
by: Zheng, Yanwei, et al.
Published: (2024)
by: Zheng, Yanwei, et al.
Published: (2024)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task Data
by: Cai, Xiongyi, et al.
Published: (2025)
by: Cai, Xiongyi, et al.
Published: (2025)
Schrödinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
by: Zantout, Nader, et al.
Published: (2025)
by: Zantout, Nader, et al.
Published: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
by: Yang, Ruihan, et al.
Published: (2025)
by: Yang, Ruihan, et al.
Published: (2025)
Diff-IP2D: Diffusion-Based Hand-Object Interaction Prediction on Egocentric Videos
by: Ma, Junyi, et al.
Published: (2024)
by: Ma, Junyi, et al.
Published: (2024)
CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning
by: Gan, Rui, et al.
Published: (2026)
by: Gan, Rui, et al.
Published: (2026)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
ZeST: an LLM-based Zero-Shot Traversability Navigation for Unknown Environments
by: Gummadi, Shreya, et al.
Published: (2025)
by: Gummadi, Shreya, et al.
Published: (2025)
Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion Descriptors
by: Tsagkas, Nikolaos, et al.
Published: (2024)
by: Tsagkas, Nikolaos, et al.
Published: (2024)
REST: Receding Horizon Explorative Steiner Tree for Zero-Shot Object-Goal Navigation
by: Xiao, Shuqi, et al.
Published: (2026)
by: Xiao, Shuqi, et al.
Published: (2026)
Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera
by: Guo, Yuliang, et al.
Published: (2025)
by: Guo, Yuliang, et al.
Published: (2025)
GVDepth: Zero-Shot Monocular Depth Estimation for Ground Vehicles based on Probabilistic Cue Fusion
by: Koledić, Karlo, et al.
Published: (2024)
by: Koledić, Karlo, et al.
Published: (2024)
Synthesizing the Kill Chain: A Zero-Shot Framework for Target Verification and Tactical Reasoning on the Edge
by: Barkley, Jesse, et al.
Published: (2026)
by: Barkley, Jesse, et al.
Published: (2026)
H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
by: Ci, Hai, et al.
Published: (2025)
by: Ci, Hai, et al.
Published: (2025)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
Gaze-Guided 3D Hand Motion Prediction for Detecting Intent in Egocentric Grasping Tasks
by: He, Yufei, et al.
Published: (2025)
by: He, Yufei, et al.
Published: (2025)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
by: Patel, Alkesh, et al.
Published: (2025)
by: Patel, Alkesh, et al.
Published: (2025)
FLAME: Learning to Navigate with Multimodal LLM in Urban Environments
by: Xu, Yunzhe, et al.
Published: (2024)
by: Xu, Yunzhe, et al.
Published: (2024)
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
by: Wang, Yunheng, et al.
Published: (2025)
by: Wang, Yunheng, et al.
Published: (2025)
Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
by: Nie, Chang, et al.
Published: (2026)
by: Nie, Chang, et al.
Published: (2026)
Object-Shot Enhanced Grounding Network for Egocentric Video
by: Feng, Yisen, et al.
Published: (2025)
by: Feng, Yisen, et al.
Published: (2025)
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
by: Yajima, Masaru, et al.
Published: (2025)
by: Yajima, Masaru, et al.
Published: (2025)
Hand-Object Interaction Pretraining from Videos
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
by: Goswami, Raktim Gautam, et al.
Published: (2025)
by: Goswami, Raktim Gautam, et al.
Published: (2025)
OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation
by: Jawaid, Ahad, et al.
Published: (2025)
by: Jawaid, Ahad, et al.
Published: (2025)
VITAL: Interactive Few-Shot Imitation Learning via Visual Human-in-the-Loop Corrections
by: Kasaei, Hamidreza, et al.
Published: (2024)
by: Kasaei, Hamidreza, et al.
Published: (2024)
Similar Items
-
EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos
by: Ma, Junyi, et al.
Published: (2025) -
Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
by: Ma, Junyi, et al.
Published: (2025) -
Robot Learning from Human Videos: A Survey
by: Ma, Junyi, et al.
Published: (2026) -
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
by: Wang, Qineng, et al.
Published: (2025) -
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
by: Xiao, Junbin, et al.
Published: (2026)