MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Jinge, Zhang, Xiaoyan, Vinod, Gautham, Raghavan, Siddeshwar, He, Jiangpeng, Zhu, Fengqing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908857281806336
author Ma, Jinge
Zhang, Xiaoyan
Vinod, Gautham
Raghavan, Siddeshwar
He, Jiangpeng
Zhu, Fengqing
author_facet Ma, Jinge
Zhang, Xiaoyan
Vinod, Gautham
Raghavan, Siddeshwar
He, Jiangpeng
Zhu, Fengqing
contents Food portion estimation is crucial for monitoring health and tracking dietary intake. Image-based dietary assessment, which involves analyzing eating occasion images using computer vision techniques, is increasingly replacing traditional methods such as 24-hour recalls. However, accurately estimating the nutritional content from images remains challenging due to the loss of 3D information when projecting to the 2D image plane. Existing portion estimation methods are challenging to deploy in real-world scenarios due to their reliance on specific requirements, such as physical reference objects, high-quality depth information, or multi-view images and videos. In this paper, we introduce MFP3D, a new framework for accurate food portion estimation using only a single monocular image. Specifically, MFP3D consists of three key modules: (1) a 3D Reconstruction Module that generates a 3D point cloud representation of the food from the 2D image, (2) a Feature Extraction Module that extracts and concatenates features from both the 3D point cloud and the 2D RGB image, and (3) a Portion Regression Module that employs a deep regression model to estimate the food's volume and energy content based on the extracted features. Our MFP3D is evaluated on MetaFood3D dataset, demonstrating its significant improvement in accurate portion estimation over existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10492
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds
Ma, Jinge
Zhang, Xiaoyan
Vinod, Gautham
Raghavan, Siddeshwar
He, Jiangpeng
Zhu, Fengqing
Computer Vision and Pattern Recognition
Image and Video Processing
Food portion estimation is crucial for monitoring health and tracking dietary intake. Image-based dietary assessment, which involves analyzing eating occasion images using computer vision techniques, is increasingly replacing traditional methods such as 24-hour recalls. However, accurately estimating the nutritional content from images remains challenging due to the loss of 3D information when projecting to the 2D image plane. Existing portion estimation methods are challenging to deploy in real-world scenarios due to their reliance on specific requirements, such as physical reference objects, high-quality depth information, or multi-view images and videos. In this paper, we introduce MFP3D, a new framework for accurate food portion estimation using only a single monocular image. Specifically, MFP3D consists of three key modules: (1) a 3D Reconstruction Module that generates a 3D point cloud representation of the food from the 2D image, (2) a Feature Extraction Module that extracts and concatenates features from both the 3D point cloud and the 2D RGB image, and (3) a Portion Regression Module that employs a deep regression model to estimate the food's volume and energy content based on the extracted features. Our MFP3D is evaluated on MetaFood3D dataset, demonstrating its significant improvement in accurate portion estimation over existing methods.
title MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2411.10492