Object-Shot Enhanced Grounding Network for Egocentric Video
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yisen, Zhang, Haoyu, Liu, Meng, Guan, Weili, Nie, Liqiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
by: Zhang, Haoyu, et al.
Published: (2023)
by: Zhang, Haoyu, et al.
Published: (2023)
VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026
by: Chu, Qiaohui, et al.
Published: (2026)
by: Chu, Qiaohui, et al.
Published: (2026)
Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
by: Chu, Qiaohui, et al.
Published: (2025)
by: Chu, Qiaohui, et al.
Published: (2025)
OSGNet @ Ego4D Episodic Memory Challenge 2025
by: Feng, Yisen, et al.
Published: (2025)
by: Feng, Yisen, et al.
Published: (2025)
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
by: Chu, Qiaohui, et al.
Published: (2025)
by: Chu, Qiaohui, et al.
Published: (2025)
HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
ObjectNLQ @ Ego4D Episodic Memory Challenge 2024
by: Feng, Yisen, et al.
Published: (2024)
by: Feng, Yisen, et al.
Published: (2024)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
HCQA @ Ego4D EgoSchema Challenge 2024
by: Zhang, Haoyu, et al.
Published: (2024)
by: Zhang, Haoyu, et al.
Published: (2024)
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
by: Chu, Qiaohui, et al.
Published: (2026)
by: Chu, Qiaohui, et al.
Published: (2026)
OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
by: Feng, Yisen, et al.
Published: (2026)
by: Feng, Yisen, et al.
Published: (2026)
Detecting Deepfakes via Hamiltonian Dynamics
by: Cheng, Harry, et al.
Published: (2026)
by: Cheng, Harry, et al.
Published: (2026)
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
by: Lall, Vishakha, et al.
Published: (2025)
by: Lall, Vishakha, et al.
Published: (2025)
Zero-Shot Temporal Interaction Localization for Egocentric Videos
by: Zhang, Erhang, et al.
Published: (2025)
by: Zhang, Erhang, et al.
Published: (2025)
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
by: Xiao, Junbin, et al.
Published: (2026)
by: Xiao, Junbin, et al.
Published: (2026)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
by: Feng, Weixi, et al.
Published: (2025)
by: Feng, Weixi, et al.
Published: (2025)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
by: Xu, Qi'ao, et al.
Published: (2025)
by: Xu, Qi'ao, et al.
Published: (2025)
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records
by: Lyu, Yibo, et al.
Published: (2026)
by: Lyu, Yibo, et al.
Published: (2026)
Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge
by: Zhang, Jinrong, et al.
Published: (2026)
by: Zhang, Jinrong, et al.
Published: (2026)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
by: Wen, Haokun, et al.
Published: (2026)
by: Wen, Haokun, et al.
Published: (2026)
Visual Self-paced Iterative Learning for Unsupervised Temporal Action Localization
by: Hu, Yupeng, et al.
Published: (2023)
by: Hu, Yupeng, et al.
Published: (2023)
Mastering Negation: Boosting Grounding Models via Grouped Opposition-Based Learning
by: Yang, Zesheng, et al.
Published: (2026)
by: Yang, Zesheng, et al.
Published: (2026)
Grounding 3D Scene Affordance From Egocentric Interactions
by: Liu, Cuiyu, et al.
Published: (2024)
by: Liu, Cuiyu, et al.
Published: (2024)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
by: Fu, Hongming, et al.
Published: (2026)
by: Fu, Hongming, et al.
Published: (2026)
ShoulderShot: Generating Over-the-Shoulder Dialogue Videos
by: Zhang, Yuang, et al.
Published: (2025)
by: Zhang, Yuang, et al.
Published: (2025)
EgoVLM: Policy Optimization for Egocentric Video Understanding
by: Vinod, Ashwin, et al.
Published: (2025)
by: Vinod, Ashwin, et al.
Published: (2025)
Object Aware Egocentric Online Action Detection
by: An, Joungbin, et al.
Published: (2024)
by: An, Joungbin, et al.
Published: (2024)
Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics
by: Tse, Tze Ho Elden, et al.
Published: (2025)
by: Tse, Tze Ho Elden, et al.
Published: (2025)
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
by: Guo, Chaohong, et al.
Published: (2025)
by: Guo, Chaohong, et al.
Published: (2025)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation
by: He, Xusheng, et al.
Published: (2026)
by: He, Xusheng, et al.
Published: (2026)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2025)
by: Seth, Ashish, et al.
Published: (2025)
Comparing Learning Paradigms for Egocentric Video Summarization
by: Wen, Daniel
Published: (2025)
by: Wen, Daniel
Published: (2025)
Exploring Audio Hallucination in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos
by: Banerjee, Prithviraj, et al.
Published: (2024)
by: Banerjee, Prithviraj, et al.
Published: (2024)
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
by: Lv, Guannan, et al.
Published: (2026)
by: Lv, Guannan, et al.
Published: (2026)
EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026
by: Fu, Zhiheng, et al.
Published: (2026)
by: Fu, Zhiheng, et al.
Published: (2026)
Similar Items
-
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
by: Zhang, Haoyu, et al.
Published: (2023) -
VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026
by: Chu, Qiaohui, et al.
Published: (2026) -
Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
by: Chu, Qiaohui, et al.
Published: (2025) -
OSGNet @ Ego4D Episodic Memory Challenge 2025
by: Feng, Yisen, et al.
Published: (2025) -
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
by: Chu, Qiaohui, et al.
Published: (2025)