FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917638446252032 |
|---|---|
| author | Yin, Yuehao Qi, Huiyan Zhu, Bin Chen, Jingjing Jiang, Yu-Gang Ngo, Chong-Wah |
| author_facet | Yin, Yuehao Qi, Huiyan Zhu, Bin Chen, Jingjing Jiang, Yu-Gang Ngo, Chong-Wah |
| contents | Large Multi-modal Models (LMMs) have made impressive progress in many vision-language tasks. Nevertheless, the performance of general LMMs in specific domains is still far from satisfactory. This paper proposes FoodLMM, a versatile food assistant based on LMMs with various capabilities, including food recognition, ingredient recognition, recipe generation, nutrition estimation, food segmentation and multi-round conversation. To facilitate FoodLMM to deal with tasks beyond pure text output, we introduce a series of novel task-specific tokens and heads, enabling the model to predict food nutritional values and multiple segmentation masks. We adopt a two-stage training strategy. In the first stage, we utilize multiple public food benchmarks for multi-task learning by leveraging the instruct-following paradigm. In the second stage, we construct a multi-round conversation dataset and a reasoning segmentation dataset to fine-tune the model, enabling it to conduct professional dialogues and generate segmentation masks based on complex reasoning in the food domain. Our fine-tuned FoodLMM achieves state-of-the-art results across several food benchmarks. We will make our code, models and datasets publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_14991 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | FoodLMM: A Versatile Food Assistant using Large Multi-modal Model Yin, Yuehao Qi, Huiyan Zhu, Bin Chen, Jingjing Jiang, Yu-Gang Ngo, Chong-Wah Computer Vision and Pattern Recognition Large Multi-modal Models (LMMs) have made impressive progress in many vision-language tasks. Nevertheless, the performance of general LMMs in specific domains is still far from satisfactory. This paper proposes FoodLMM, a versatile food assistant based on LMMs with various capabilities, including food recognition, ingredient recognition, recipe generation, nutrition estimation, food segmentation and multi-round conversation. To facilitate FoodLMM to deal with tasks beyond pure text output, we introduce a series of novel task-specific tokens and heads, enabling the model to predict food nutritional values and multiple segmentation masks. We adopt a two-stage training strategy. In the first stage, we utilize multiple public food benchmarks for multi-task learning by leveraging the instruct-following paradigm. In the second stage, we construct a multi-round conversation dataset and a reasoning segmentation dataset to fine-tune the model, enabling it to conduct professional dialogues and generate segmentation masks based on complex reasoning in the food domain. Our fine-tuned FoodLMM achieves state-of-the-art results across several food benchmarks. We will make our code, models and datasets publicly available. |
| title | FoodLMM: A Versatile Food Assistant using Large Multi-modal Model |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2312.14991 |