From Videos to Conversations: Egocentric Instructions for Task Assistance
Fuente:
arXiv
Saved in:
| Main Authors: | Aggarwal, Lavisha, Bahirwani, Vikas, Colaco, Andrea |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
by: Aggarwal, Lavisha, et al.
Published: (2025)
by: Aggarwal, Lavisha, et al.
Published: (2025)
ESAM++: Efficient Online 3D Perception on the Edge
by: Liu, Qin, et al.
Published: (2026)
by: Liu, Qin, et al.
Published: (2026)
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion
by: Tian, Junjiao, et al.
Published: (2023)
by: Tian, Junjiao, et al.
Published: (2023)
Identification of Conversation Partners from Egocentric Video
by: Dorszewski, Tobias, et al.
Published: (2024)
by: Dorszewski, Tobias, et al.
Published: (2024)
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
by: Toschi, Federico, et al.
Published: (2026)
by: Toschi, Federico, et al.
Published: (2026)
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
by: Jia, Wenqi, et al.
Published: (2023)
by: Jia, Wenqi, et al.
Published: (2023)
Simultaneous Localization and Affordance Prediction of Tasks from Egocentric Video
by: Chavis, Zachary, et al.
Published: (2024)
by: Chavis, Zachary, et al.
Published: (2024)
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2025)
by: Seminara, Luigi, et al.
Published: (2025)
Hier-EgoPack: Hierarchical Egocentric Video Understanding with Diverse Task Perspectives
by: Peirone, Simone Alberto, et al.
Published: (2025)
by: Peirone, Simone Alberto, et al.
Published: (2025)
HEADS-UP: Head-Mounted Egocentric Dataset for Trajectory Prediction in Blind Assistance Systems
by: Haghighi, Yasaman, et al.
Published: (2024)
by: Haghighi, Yasaman, et al.
Published: (2024)
Exploring Audio Hallucination in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
by: Sakai, Yuki, et al.
Published: (2025)
by: Sakai, Yuki, et al.
Published: (2025)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
by: Xiao, Junbin, et al.
Published: (2025)
by: Xiao, Junbin, et al.
Published: (2025)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
by: Liu, Ruyang, et al.
Published: (2023)
by: Liu, Ruyang, et al.
Published: (2023)
EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses
by: Paruchuri, Akshay, et al.
Published: (2025)
by: Paruchuri, Akshay, et al.
Published: (2025)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
by: Ohkawa, Takehiko, et al.
Published: (2023)
by: Ohkawa, Takehiko, et al.
Published: (2023)
Retrieval-Augmented Egocentric Video Captioning
by: Xu, Jilan, et al.
Published: (2024)
by: Xu, Jilan, et al.
Published: (2024)
Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos
by: Tschernezki, Vadim, et al.
Published: (2025)
by: Tschernezki, Vadim, et al.
Published: (2025)
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2024)
by: Seminara, Luigi, et al.
Published: (2024)
FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video
by: Camiletto, Andrea Boscolo, et al.
Published: (2025)
by: Camiletto, Andrea Boscolo, et al.
Published: (2025)
Spatial-Conditioned Reasoning in Long-Egocentric Videos
by: Tribble, James, et al.
Published: (2026)
by: Tribble, James, et al.
Published: (2026)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
EgoTV: Egocentric Task Verification from Natural Language Task Descriptions
by: Hazra, Rishi, et al.
Published: (2023)
by: Hazra, Rishi, et al.
Published: (2023)
A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives
by: Peirone, Simone Alberto, et al.
Published: (2024)
by: Peirone, Simone Alberto, et al.
Published: (2024)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
by: Xu, Qi'ao, et al.
Published: (2025)
by: Xu, Qi'ao, et al.
Published: (2025)
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
by: Wang, Yanan, et al.
Published: (2025)
by: Wang, Yanan, et al.
Published: (2025)
SurfaceXR: Fusing Smartwatch IMUs and Egocentric Hand Pose for Seamless Surface Interactions
by: Xu, Vasco, et al.
Published: (2026)
by: Xu, Vasco, et al.
Published: (2026)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
TEXT2TASTE: A Versatile Egocentric Vision System for Intelligent Reading Assistance Using Large Language Model
by: Mucha, Wiktor, et al.
Published: (2024)
by: Mucha, Wiktor, et al.
Published: (2024)
MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
What to Do Next? Memorizing skills from Egocentric Instructional Video
by: Bi, Jing, et al.
Published: (2025)
by: Bi, Jing, et al.
Published: (2025)
Motion Focus Recognition in Fast-Moving Egocentric Video
by: Hong, Si-En, et al.
Published: (2026)
by: Hong, Si-En, et al.
Published: (2026)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
by: Zhu, Bingwen, et al.
Published: (2026)
by: Zhu, Bingwen, et al.
Published: (2026)
Detecting Precise Hand Touch Moments in Egocentric Video
by: Nguyen, Huy Anh, et al.
Published: (2026)
by: Nguyen, Huy Anh, et al.
Published: (2026)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
by: Xu, Zhiyang, et al.
Published: (2026)
by: Xu, Zhiyang, et al.
Published: (2026)
EgoPoints: Advancing Point Tracking for Egocentric Videos
by: Darkhalil, Ahmad, et al.
Published: (2024)
by: Darkhalil, Ahmad, et al.
Published: (2024)
EgoAVU: Egocentric Audio-Visual Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
EgoQR: Efficient QR Code Reading in Egocentric Settings
by: Moslehpour, Mohsen, et al.
Published: (2024)
by: Moslehpour, Mohsen, et al.
Published: (2024)
Similar Items
-
Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
by: Aggarwal, Lavisha, et al.
Published: (2025) -
ESAM++: Efficient Online 3D Perception on the Edge
by: Liu, Qin, et al.
Published: (2026) -
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025) -
Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion
by: Tian, Junjiao, et al.
Published: (2023) -
Identification of Conversation Partners from Egocentric Video
by: Dorszewski, Tobias, et al.
Published: (2024)