Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Aggarwal, Lavisha, Bahirwani, Vikas, Li, Lin, Colaco, Andrea
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908490873700352
author Aggarwal, Lavisha
Bahirwani, Vikas
Li, Lin
Colaco, Andrea
author_facet Aggarwal, Lavisha
Bahirwani, Vikas
Li, Lin
Colaco, Andrea
contents Many everyday tasks ranging from fixing appliances, cooking recipes to car maintenance require expert knowledge, especially when tasks are complex and multi-step. Despite growing interest in AI agents, there is a scarcity of dialogue-video datasets grounded for real world task assistance. In this paper, we propose a simple yet effective approach that transforms single-person instructional videos into task-guidance two-person dialogues, aligned with fine grained steps and video-clips. Our fully automatic approach, powered by large language models, offers an efficient alternative to the substantial cost and effort required for human-assisted data collection. Using this technique, we build HowToDIV, a large-scale dataset containing 507 conversations, 6636 question-answer pairs and 24 hours of videoclips across diverse tasks in cooking, mechanics, and planting. Each session includes multi-turn conversation where an expert teaches a novice user how to perform a task step by step, while observing user's surrounding through a camera and microphone equipped wearable device. We establish the baseline benchmark performance on HowToDIV dataset through Gemma-3 model for future research on this new task of dialogues for procedural-task assistance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11192
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
Aggarwal, Lavisha
Bahirwani, Vikas
Li, Lin
Colaco, Andrea
Computer Vision and Pattern Recognition
Many everyday tasks ranging from fixing appliances, cooking recipes to car maintenance require expert knowledge, especially when tasks are complex and multi-step. Despite growing interest in AI agents, there is a scarcity of dialogue-video datasets grounded for real world task assistance. In this paper, we propose a simple yet effective approach that transforms single-person instructional videos into task-guidance two-person dialogues, aligned with fine grained steps and video-clips. Our fully automatic approach, powered by large language models, offers an efficient alternative to the substantial cost and effort required for human-assisted data collection. Using this technique, we build HowToDIV, a large-scale dataset containing 507 conversations, 6636 question-answer pairs and 24 hours of videoclips across diverse tasks in cooking, mechanics, and planting. Each session includes multi-turn conversation where an expert teaches a novice user how to perform a task step by step, while observing user's surrounding through a camera and microphone equipped wearable device. We establish the baseline benchmark performance on HowToDIV dataset through Gemma-3 model for future research on this new task of dialogues for procedural-task assistance.
title Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.11192