Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Kundu, Sanjoy, Trehan, Shubham, Aakur, Sathyanarayanan N. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
by: Kundu, Sanjoy, et al.
Published: (2024)
by: Kundu, Sanjoy, et al.
Published: (2024)
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025)
by: Kundu, Sanjoy, et al.
Published: (2025)
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025)
by: Kundu, Sanjoy, et al.
Published: (2025)
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
by: Trehan, Shubham, et al.
Published: (2024)
by: Trehan, Shubham, et al.
Published: (2024)
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
by: Vellamcheti, Shanmukha, et al.
Published: (2025)
by: Vellamcheti, Shanmukha, et al.
Published: (2025)
FSP-DETR: Few-Shot Prototypical Parasitic Ova Detection
by: Trehan, Shubham, et al.
Published: (2025)
by: Trehan, Shubham, et al.
Published: (2025)
EASE: Embodied Active Event Perception via Self-Supervised Energy Minimization
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
STaTS: Structure-Aware Temporal Sequence Summarization via Statistical Window Merging
by: Bhowmick, Disharee, et al.
Published: (2025)
by: Bhowmick, Disharee, et al.
Published: (2025)
WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos
by: Ye, Yufei, et al.
Published: (2026)
by: Ye, Yufei, et al.
Published: (2026)
Capturing Temporal Components for Time Series Classification
by: Vavilthota, Venkata Ragavendra, et al.
Published: (2024)
by: Vavilthota, Venkata Ragavendra, et al.
Published: (2024)
CVT-Bench: Counterfactual Viewpoint Transformations Reveal Unstable Spatial Representations in Multimodal LLMs
by: Vellamcheti, Shanmukha, et al.
Published: (2026)
by: Vellamcheti, Shanmukha, et al.
Published: (2026)
Improving Open-World Object Localization by Discovering Background
by: Singh, Ashish, et al.
Published: (2025)
by: Singh, Ashish, et al.
Published: (2025)
Object-Shot Enhanced Grounding Network for Egocentric Video
by: Feng, Yisen, et al.
Published: (2025)
by: Feng, Yisen, et al.
Published: (2025)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
by: Liu, Huabin, et al.
Published: (2025)
by: Liu, Huabin, et al.
Published: (2025)
A Study of Commonsense Reasoning over Visual Object Properties
by: Kolari, Abhishek, et al.
Published: (2025)
by: Kolari, Abhishek, et al.
Published: (2025)
LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
by: Gao, Quankai, et al.
Published: (2026)
by: Gao, Quankai, et al.
Published: (2026)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
by: Yun, Heeseung, et al.
Published: (2024)
by: Yun, Heeseung, et al.
Published: (2024)
Visual Intention Grounding for Egocentric Assistants
by: Sun, Pengzhan, et al.
Published: (2025)
by: Sun, Pengzhan, et al.
Published: (2025)
Augmented Commonsense Knowledge for Remote Object Grounding
by: Mohammadi, Bahram, et al.
Published: (2024)
by: Mohammadi, Bahram, et al.
Published: (2024)
Causal Debiasing for Visual Commonsense Reasoning
by: Zou, Jiayi, et al.
Published: (2025)
by: Zou, Jiayi, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
by: Huang, Wenyuan, et al.
Published: (2025)
by: Huang, Wenyuan, et al.
Published: (2025)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
Fine-grained Spatiotemporal Grounding on Egocentric Videos
by: Liang, Shuo, et al.
Published: (2025)
by: Liang, Shuo, et al.
Published: (2025)
Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
by: Wei, Jiude, et al.
Published: (2025)
by: Wei, Jiude, et al.
Published: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
Interaction Region Visual Transformer for Egocentric Action Anticipation
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
by: Plizzari, Chiara, et al.
Published: (2024)
by: Plizzari, Chiara, et al.
Published: (2024)
Spatial-Conditioned Reasoning in Long-Egocentric Videos
by: Tribble, James, et al.
Published: (2026)
by: Tribble, James, et al.
Published: (2026)
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024)
by: An, Joungbin, et al.
Published: (2024)
Object-centric Video Question Answering with Visual Grounding and Referring
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection
by: Zhuang, Weijun, et al.
Published: (2025)
by: Zhuang, Weijun, et al.
Published: (2025)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
by: Zhang, Mingfang, et al.
Published: (2024)
by: Zhang, Mingfang, et al.
Published: (2024)
Similar Items
-
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
by: Kundu, Sanjoy, et al.
Published: (2024) -
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025) -
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025) -
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
by: Trehan, Shubham, et al.
Published: (2024) -
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
by: Vellamcheti, Shanmukha, et al.
Published: (2025)