CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ruoxuan, Wen, Bin, Xie, Hongxia, Yao, Yi, Zuo, Songhan, Jiang-Lin, Jian-Yu, Shuai, Hong-Han, Cheng, Wen-Huang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911303502659584
author Zhang, Ruoxuan
Wen, Bin
Xie, Hongxia
Yao, Yi
Zuo, Songhan
Jiang-Lin, Jian-Yu
Shuai, Hong-Han
Cheng, Wen-Huang
author_facet Zhang, Ruoxuan
Wen, Bin
Xie, Hongxia
Yao, Yi
Zuo, Songhan
Jiang-Lin, Jian-Yu
Shuai, Hong-Han
Cheng, Wen-Huang
contents Cooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusion models have shown strong capabilities in text-to-image generation, they struggle to handle structured multi-step scenarios like recipe illustration. Additionally, current recipe illustration methods are unable to adjust to the natural variability in recipe length, generating a fixed number of images regardless of the actual instructions structure. To address these limitations, we present CookAnything, a flexible and consistent diffusion-based framework that generates coherent, semantically distinct image sequences from textual cooking instructions of arbitrary length. The framework introduces three key components: (1) Step-wise Regional Control (SRC), which aligns textual steps with corresponding image regions within a single denoising process; (2) Flexible RoPE, a step-aware positional encoding mechanism that enhances both temporal coherence and spatial diversity; and (3) Cross-Step Consistency Control (CSCC), which maintains fine-grained ingredient consistency across steps. Experimental results on recipe illustration benchmarks show that CookAnything performs better than existing methods in training-based and training-free settings. The proposed framework supports scalable, high-quality visual synthesis of complex multi-step instructions and holds significant potential for broad applications in instructional media, and procedural content creation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03540
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
Zhang, Ruoxuan
Wen, Bin
Xie, Hongxia
Yao, Yi
Zuo, Songhan
Jiang-Lin, Jian-Yu
Shuai, Hong-Han
Cheng, Wen-Huang
Computer Vision and Pattern Recognition
Artificial Intelligence
Cooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusion models have shown strong capabilities in text-to-image generation, they struggle to handle structured multi-step scenarios like recipe illustration. Additionally, current recipe illustration methods are unable to adjust to the natural variability in recipe length, generating a fixed number of images regardless of the actual instructions structure. To address these limitations, we present CookAnything, a flexible and consistent diffusion-based framework that generates coherent, semantically distinct image sequences from textual cooking instructions of arbitrary length. The framework introduces three key components: (1) Step-wise Regional Control (SRC), which aligns textual steps with corresponding image regions within a single denoising process; (2) Flexible RoPE, a step-aware positional encoding mechanism that enhances both temporal coherence and spatial diversity; and (3) Cross-Step Consistency Control (CSCC), which maintains fine-grained ingredient consistency across steps. Experimental results on recipe illustration benchmarks show that CookAnything performs better than existing methods in training-based and training-free settings. The proposed framework supports scalable, high-quality visual synthesis of complex multi-step instructions and holds significant potential for broad applications in instructional media, and procedural content creation.
title CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.03540