EcomMMMU: Strategic Utilization of Visuals for Robust Multimodal E-commerce Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ling, Xinyi, Du, Hanwen, Zhu, Zhihui, Ning, Xia
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915613188816896
author Ling, Xinyi
Du, Hanwen
Zhu, Zhihui
Ning, Xia
author_facet Ling, Xinyi
Du, Hanwen
Zhu, Zhihui
Ning, Xia
contents E-commerce platforms are rich in multimodal data, featuring a variety of images that depict product details. However, this raises an important question: do these images always enhance product understanding, or can they sometimes introduce redundancy or degrade performance? Existing datasets are limited in both scale and design, making it difficult to systematically examine this question. To this end, we introduce EcomMMMU, an e-commerce multimodal multitask understanding dataset with 406,190 samples and 8,989,510 images. EcomMMMU is comprised of multi-image visual-language data designed with 8 essential tasks and a specialized VSS subset to benchmark the capability of multimodal large language models (MLLMs) to effectively utilize visual content. Analysis on EcomMMMU reveals that product images do not consistently improve performance and can, in some cases, degrade it. This indicates that MLLMs may struggle to effectively leverage rich visual content for e-commerce tasks. Building on these insights, we propose SUMEI, a data-driven method that strategically utilizes multiple images via predicting visual utilities before using them for downstream tasks. Comprehensive experiments demonstrate the effectiveness and robustness of SUMEI. The data and code are available through https://github.com/ninglab/EcomMMMU.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15721
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EcomMMMU: Strategic Utilization of Visuals for Robust Multimodal E-commerce Models
Ling, Xinyi
Du, Hanwen
Zhu, Zhihui
Ning, Xia
Computation and Language
Artificial Intelligence
E-commerce platforms are rich in multimodal data, featuring a variety of images that depict product details. However, this raises an important question: do these images always enhance product understanding, or can they sometimes introduce redundancy or degrade performance? Existing datasets are limited in both scale and design, making it difficult to systematically examine this question. To this end, we introduce EcomMMMU, an e-commerce multimodal multitask understanding dataset with 406,190 samples and 8,989,510 images. EcomMMMU is comprised of multi-image visual-language data designed with 8 essential tasks and a specialized VSS subset to benchmark the capability of multimodal large language models (MLLMs) to effectively utilize visual content. Analysis on EcomMMMU reveals that product images do not consistently improve performance and can, in some cases, degrade it. This indicates that MLLMs may struggle to effectively leverage rich visual content for e-commerce tasks. Building on these insights, we propose SUMEI, a data-driven method that strategically utilizes multiple images via predicting visual utilities before using them for downstream tasks. Comprehensive experiments demonstrate the effectiveness and robustness of SUMEI. The data and code are available through https://github.com/ninglab/EcomMMMU.
title EcomMMMU: Strategic Utilization of Visuals for Robust Multimodal E-commerce Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.15721