Food Image Generation on Multi-Noun Categories

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pan, Xinyue, Chen, Yuhao, He, Jiangpeng, Zhu, Fengqing
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911310805991424
author Pan, Xinyue
Chen, Yuhao
He, Jiangpeng
Zhu, Fengqing
author_facet Pan, Xinyue
Chen, Yuhao
He, Jiangpeng
Zhu, Fengqing
contents Generating realistic food images for categories with multiple nouns is surprisingly challenging. For instance, the prompt "egg noodle" may result in images that incorrectly contain both eggs and noodles as separate entities. Multi-noun food categories are common in real-world datasets and account for a large portion of entries in benchmarks such as UEC-256. These compound names often cause generative models to misinterpret the semantics, producing unintended ingredients or objects. This is due to insufficient multi-noun category related knowledge in the text encoder and misinterpretation of multi-noun relationships, leading to incorrect spatial layouts. To overcome these challenges, we propose FoCULR (Food Category Understanding and Layout Refinement) which incorporates food domain knowledge and introduces core concepts early in the generation process. Experimental results demonstrate that the integration of these techniques improves image generation performance in the food domain.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09095
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Food Image Generation on Multi-Noun Categories
Pan, Xinyue
Chen, Yuhao
He, Jiangpeng
Zhu, Fengqing
Computer Vision and Pattern Recognition
Generating realistic food images for categories with multiple nouns is surprisingly challenging. For instance, the prompt "egg noodle" may result in images that incorrectly contain both eggs and noodles as separate entities. Multi-noun food categories are common in real-world datasets and account for a large portion of entries in benchmarks such as UEC-256. These compound names often cause generative models to misinterpret the semantics, producing unintended ingredients or objects. This is due to insufficient multi-noun category related knowledge in the text encoder and misinterpretation of multi-noun relationships, leading to incorrect spatial layouts. To overcome these challenges, we propose FoCULR (Food Category Understanding and Layout Refinement) which incorporates food domain knowledge and introduces core concepts early in the generation process. Experimental results demonstrate that the integration of these techniques improves image generation performance in the food domain.
title Food Image Generation on Multi-Noun Categories
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.09095