RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909645689323520 |
|---|---|
| author | Zhang, Ruoxuan Gao, Jidong Wen, Bin Xie, Hongxia Zhang, Chenming Shuai, Hong-Han Cheng, Wen-Huang |
| author_facet | Zhang, Ruoxuan Gao, Jidong Wen, Bin Xie, Hongxia Zhang, Chenming Shuai, Hong-Han Cheng, Wen-Huang |
| contents | Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and visual content. We present RecipeGen, the first large-scale, real-world benchmark for recipe-based Text-to-Image (T2I), Image-to-Video (I2V), and Text-to-Video (T2V) generation. RecipeGen contains 26,453 recipes, 196,724 images, and 4,491 videos, covering diverse ingredients, cooking procedures, styles, and dish types. We further propose domain-specific evaluation metrics to assess ingredient fidelity and interaction modeling, benchmark representative T2I, I2V, and T2V models, and provide insights for future recipe generation models. Project page is available now. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_06733 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation Zhang, Ruoxuan Gao, Jidong Wen, Bin Xie, Hongxia Zhang, Chenming Shuai, Hong-Han Cheng, Wen-Huang Computer Vision and Pattern Recognition Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and visual content. We present RecipeGen, the first large-scale, real-world benchmark for recipe-based Text-to-Image (T2I), Image-to-Video (I2V), and Text-to-Video (T2V) generation. RecipeGen contains 26,453 recipes, 196,724 images, and 4,491 videos, covering diverse ingredients, cooking procedures, styles, and dish types. We further propose domain-specific evaluation metrics to assess ingredient fidelity and interaction modeling, benchmark representative T2I, I2V, and T2V models, and provide insights for future recipe generation models. Project page is available now. |
| title | RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.06733 |