EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sakai, Yuki, Furuta, Ryosuke, Yen, Juichun, Sato, Yoichi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2023)
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2023)
Seeking Flat Minima with Mean Teacher on Semi- and Weakly-Supervised Domain Generalization for Object Detection
von: Furuta, Ryosuke, et al.
Veröffentlicht: (2023)
von: Furuta, Ryosuke, et al.
Veröffentlicht: (2023)
Leadership Assessment in Pediatric Intensive Care Unit Team Training
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
Multi-speaker Attention Alignment for Multimodal Social Interaction
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
ActionVOS: Actions as Prompts for Video Object Segmentation
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2024)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2024)
FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
von: Yagi, Takuma, et al.
Veröffentlicht: (2024)
von: Yagi, Takuma, et al.
Veröffentlicht: (2024)
Learning Multiple Object States from Actions via Large Language Models
von: Tateno, Masatoshi, et al.
Veröffentlicht: (2024)
von: Tateno, Masatoshi, et al.
Veröffentlicht: (2024)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
von: Leonardi, Rosario, et al.
Veröffentlicht: (2026)
von: Leonardi, Rosario, et al.
Veröffentlicht: (2026)
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction
von: Li, Bate, et al.
Veröffentlicht: (2025)
von: Li, Bate, et al.
Veröffentlicht: (2025)
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
von: Lin, Nie, et al.
Veröffentlicht: (2025)
von: Lin, Nie, et al.
Veröffentlicht: (2025)
AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
von: Banno, Tatsuro, et al.
Veröffentlicht: (2025)
von: Banno, Tatsuro, et al.
Veröffentlicht: (2025)
EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
von: Li, Runjia, et al.
Veröffentlicht: (2025)
von: Li, Runjia, et al.
Veröffentlicht: (2025)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
von: Hummel, Thomas, et al.
Veröffentlicht: (2024)
von: Hummel, Thomas, et al.
Veröffentlicht: (2024)
Egocentric Gaze Estimation via Neck-Mounted Camera
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025)
von: Plizzari, Chiara, et al.
Veröffentlicht: (2025)
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
von: Zhang, Mingfang, et al.
Veröffentlicht: (2024)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2024)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
Map-Mono-Ego: Map-Grounded Global Human Pose Estimation from Monocular Egocentric Video
von: Deguchi, Hiroyuki, et al.
Veröffentlicht: (2026)
von: Deguchi, Hiroyuki, et al.
Veröffentlicht: (2026)
EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos
von: Ma, Junyi, et al.
Veröffentlicht: (2025)
von: Ma, Junyi, et al.
Veröffentlicht: (2025)
Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
von: Aggarwal, Lavisha, et al.
Veröffentlicht: (2025)
von: Aggarwal, Lavisha, et al.
Veröffentlicht: (2025)
Ego-1K -- A Large-Scale Multiview Video Dataset for Egocentric Vision
von: Lee, Jae Yong, et al.
Veröffentlicht: (2026)
von: Lee, Jae Yong, et al.
Veröffentlicht: (2026)
EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
Inference-time Trajectory Optimization for Manga Image Editing
von: Furuta, Ryosuke
Veröffentlicht: (2026)
von: Furuta, Ryosuke
Veröffentlicht: (2026)
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
von: Yan, Jiaqi, et al.
Veröffentlicht: (2025)
von: Yan, Jiaqi, et al.
Veröffentlicht: (2025)
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
von: Lin, Nie, et al.
Veröffentlicht: (2024)
von: Lin, Nie, et al.
Veröffentlicht: (2024)
Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
von: Suzuki, Naru, et al.
Veröffentlicht: (2025)
von: Suzuki, Naru, et al.
Veröffentlicht: (2025)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)
EgoPoints: Advancing Point Tracking for Egocentric Videos
von: Darkhalil, Ahmad, et al.
Veröffentlicht: (2024)
von: Darkhalil, Ahmad, et al.
Veröffentlicht: (2024)
EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset
von: John, Ronan, et al.
Veröffentlicht: (2025)
von: John, Ronan, et al.
Veröffentlicht: (2025)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
von: Mazzamuto, Michele, et al.
Veröffentlicht: (2024)
von: Mazzamuto, Michele, et al.
Veröffentlicht: (2024)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
von: Kulkarni, Yogesh, et al.
Veröffentlicht: (2025)
von: Kulkarni, Yogesh, et al.
Veröffentlicht: (2025)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
von: Pei, Baoqi, et al.
Veröffentlicht: (2024)
EgoLCD: Egocentric Video Generation with Long Context Diffusion
von: Zhang, Liuzhou, et al.
Veröffentlicht: (2025)
von: Zhang, Liuzhou, et al.
Veröffentlicht: (2025)
EgoGraph: Temporal Knowledge Graph for Egocentric Video Understanding
von: Sun, Shitong, et al.
Veröffentlicht: (2026)
von: Sun, Shitong, et al.
Veröffentlicht: (2026)
Instruct-Imagen: Image Generation with Multi-modal Instruction
von: Hu, Hexiang, et al.
Veröffentlicht: (2024)
von: Hu, Hexiang, et al.
Veröffentlicht: (2024)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2023) -
Seeking Flat Minima with Mean Teacher on Semi- and Weakly-Supervised Domain Generalization for Object Detection
von: Furuta, Ryosuke, et al.
Veröffentlicht: (2023) -
Leadership Assessment in Pediatric Intensive Care Unit Team Training
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025) -
Multi-speaker Attention Alignment for Multimodal Social Interaction
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025) -
ActionVOS: Actions as Prompts for Video Object Segmentation
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2024)