Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mittal, Prateek, Goyal, Puneet, Chauhan, Joohi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913975178887168
author Mittal, Prateek
Goyal, Puneet
Chauhan, Joohi
author_facet Mittal, Prateek
Goyal, Puneet
Chauhan, Joohi
contents This study introduces a novel multimodal food recognition framework that effectively combines visual and textual modalities to enhance classification accuracy and robustness. The proposed approach employs a dynamic multimodal fusion strategy that adaptively integrates features from unimodal visual inputs and complementary textual metadata. This fusion mechanism is designed to maximize the use of informative content, while mitigating the adverse impact of missing or inconsistent modality data. The framework was rigorously evaluated on the UPMC Food-101 dataset and achieved unimodal classification accuracies of 73.60% for images and 88.84% for text. When both modalities were fused, the model achieved an accuracy of 97.84%, outperforming several state-of-the-art methods. Extensive experimental analysis demonstrated the robustness, adaptability, and computational efficiency of the proposed settings, highlighting its practical applicability to real-world multimodal food-recognition scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2308_02562
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification
Mittal, Prateek
Goyal, Puneet
Chauhan, Joohi
Computer Vision and Pattern Recognition
Artificial Intelligence
Computers and Society
Machine Learning
This study introduces a novel multimodal food recognition framework that effectively combines visual and textual modalities to enhance classification accuracy and robustness. The proposed approach employs a dynamic multimodal fusion strategy that adaptively integrates features from unimodal visual inputs and complementary textual metadata. This fusion mechanism is designed to maximize the use of informative content, while mitigating the adverse impact of missing or inconsistent modality data. The framework was rigorously evaluated on the UPMC Food-101 dataset and achieved unimodal classification accuracies of 73.60% for images and 88.84% for text. When both modalities were fused, the model achieved an accuracy of 97.84%, outperforming several state-of-the-art methods. Extensive experimental analysis demonstrated the robustness, adaptability, and computational efficiency of the proposed settings, highlighting its practical applicability to real-world multimodal food-recognition scenarios.
title Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2308.02562