RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ruoxuan, Gao, Jidong, Wen, Bin, Xie, Hongxia, Zhang, Chenming, Shuai, Hong-Han, Cheng, Wen-Huang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909645689323520
author Zhang, Ruoxuan
Gao, Jidong
Wen, Bin
Xie, Hongxia
Zhang, Chenming
Shuai, Hong-Han
Cheng, Wen-Huang
author_facet Zhang, Ruoxuan
Gao, Jidong
Wen, Bin
Xie, Hongxia
Zhang, Chenming
Shuai, Hong-Han
Cheng, Wen-Huang
contents Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and visual content. We present RecipeGen, the first large-scale, real-world benchmark for recipe-based Text-to-Image (T2I), Image-to-Video (I2V), and Text-to-Video (T2V) generation. RecipeGen contains 26,453 recipes, 196,724 images, and 4,491 videos, covering diverse ingredients, cooking procedures, styles, and dish types. We further propose domain-specific evaluation metrics to assess ingredient fidelity and interaction modeling, benchmark representative T2I, I2V, and T2V models, and provide insights for future recipe generation models. Project page is available now.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06733
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
Zhang, Ruoxuan
Gao, Jidong
Wen, Bin
Xie, Hongxia
Zhang, Chenming
Shuai, Hong-Han
Cheng, Wen-Huang
Computer Vision and Pattern Recognition
Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and visual content. We present RecipeGen, the first large-scale, real-world benchmark for recipe-based Text-to-Image (T2I), Image-to-Video (I2V), and Text-to-Video (T2V) generation. RecipeGen contains 26,453 recipes, 196,724 images, and 4,491 videos, covering diverse ingredients, cooking procedures, styles, and dish types. We further propose domain-specific evaluation metrics to assess ingredient fidelity and interaction modeling, benchmark representative T2I, I2V, and T2V models, and provide insights for future recipe generation models. Project page is available now.
title RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.06733