VisualChef: Generating Visual Aids in Cooking via Mask Inpainting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuzyk, Oleh, Li, Zuoyue, Pollefeys, Marc, Wang, Xi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916989573791744
author Kuzyk, Oleh
Li, Zuoyue
Pollefeys, Marc
Wang, Xi
author_facet Kuzyk, Oleh
Li, Zuoyue
Pollefeys, Marc
Wang, Xi
contents Cooking requires not only following instructions but also understanding, executing, and monitoring each step - a process that can be challenging without visual guidance. Although recipe images and videos offer helpful cues, they often lack consistency in focus, tools, and setup. To better support the cooking process, we introduce VisualChef, a method for generating contextual visual aids tailored to cooking scenarios. Given an initial frame and a specified action, VisualChef generates images depicting both the action's execution and the resulting appearance of the object, while preserving the initial frame's environment. Previous work aims to integrate knowledge extracted from large language models by generating detailed textual descriptions to guide image generation, which requires fine-grained visual-textual alignment and involves additional annotations. In contrast, VisualChef simplifies alignment through mask-based visual grounding. Our key insight is identifying action-relevant objects and classifying them to enable targeted modifications that reflect the intended action and outcome while maintaining a consistent environment. In addition, we propose an automated pipeline to extract high-quality initial, action, and final state frames. We evaluate VisualChef quantitatively and qualitatively on three egocentric video datasets and show its improvements over state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18569
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VisualChef: Generating Visual Aids in Cooking via Mask Inpainting
Kuzyk, Oleh
Li, Zuoyue
Pollefeys, Marc
Wang, Xi
Computer Vision and Pattern Recognition
Cooking requires not only following instructions but also understanding, executing, and monitoring each step - a process that can be challenging without visual guidance. Although recipe images and videos offer helpful cues, they often lack consistency in focus, tools, and setup. To better support the cooking process, we introduce VisualChef, a method for generating contextual visual aids tailored to cooking scenarios. Given an initial frame and a specified action, VisualChef generates images depicting both the action's execution and the resulting appearance of the object, while preserving the initial frame's environment. Previous work aims to integrate knowledge extracted from large language models by generating detailed textual descriptions to guide image generation, which requires fine-grained visual-textual alignment and involves additional annotations. In contrast, VisualChef simplifies alignment through mask-based visual grounding. Our key insight is identifying action-relevant objects and classifying them to enable targeted modifications that reflect the intended action and outcome while maintaining a consistent environment. In addition, we propose an automated pipeline to extract high-quality initial, action, and final state frames. We evaluate VisualChef quantitatively and qualitatively on three egocentric video datasets and show its improvements over state-of-the-art methods.
title VisualChef: Generating Visual Aids in Cooking via Mask Inpainting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.18569