Adapting Vision-Language Models for E-commerce Understanding at Scale

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nulli, Matteo, Orshulevich, Vladimir, Bazazo, Tala, Herold, Christian, Kozielski, Michael, Mazur, Marcin, Tuzel, Szymon, Snoek, Cees G. M., Hashemi, Seyyed Hadi, Javed, Omar, Versley, Yannick, Khadivi, Shahram
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914324762591232
author Nulli, Matteo
Orshulevich, Vladimir
Bazazo, Tala
Herold, Christian
Kozielski, Michael
Mazur, Marcin
Tuzel, Szymon
Snoek, Cees G. M.
Hashemi, Seyyed Hadi
Javed, Omar
Versley, Yannick
Khadivi, Shahram
author_facet Nulli, Matteo
Orshulevich, Vladimir
Bazazo, Tala
Herold, Christian
Kozielski, Michael
Mazur, Marcin
Tuzel, Szymon
Snoek, Cees G. M.
Hashemi, Seyyed Hadi
Javed, Omar
Versley, Yannick
Khadivi, Shahram
contents E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent modelling, yet there is no documented, well-known strategy for adapting them to the attribute-centric, multi-image, and noisy nature of e-commerce data, without sacrificing general performance. In this work, we show through a large-scale experimental study, how targeted adaptation of general VLMs can substantially improve e-commerce performance while preserving broad multimodal capabilities. Furthermore, we propose a novel extensive evaluation suite covering deep product understanding, strict instruction following, and dynamic attribute extraction.
format Preprint
id arxiv_https___arxiv_org_abs_2602_11733
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adapting Vision-Language Models for E-commerce Understanding at Scale
Nulli, Matteo
Orshulevich, Vladimir
Bazazo, Tala
Herold, Christian
Kozielski, Michael
Mazur, Marcin
Tuzel, Szymon
Snoek, Cees G. M.
Hashemi, Seyyed Hadi
Javed, Omar
Versley, Yannick
Khadivi, Shahram
Computer Vision and Pattern Recognition
Artificial Intelligence
E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent modelling, yet there is no documented, well-known strategy for adapting them to the attribute-centric, multi-image, and noisy nature of e-commerce data, without sacrificing general performance. In this work, we show through a large-scale experimental study, how targeted adaptation of general VLMs can substantially improve e-commerce performance while preserving broad multimodal capabilities. Furthermore, we propose a novel extensive evaluation suite covering deep product understanding, strict instruction following, and dynamic attribute extraction.
title Adapting Vision-Language Models for E-commerce Understanding at Scale
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.11733