Adapting Vision-Language Models for E-commerce Understanding at Scale

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nulli, Matteo, Orshulevich, Vladimir, Bazazo, Tala, Herold, Christian, Kozielski, Michael, Mazur, Marcin, Tuzel, Szymon, Snoek, Cees G. M., Hashemi, Seyyed Hadi, Javed, Omar, Versley, Yannick, Khadivi, Shahram
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914324762591232
author Nulli, Matteo
Orshulevich, Vladimir
Bazazo, Tala
Herold, Christian
Kozielski, Michael
Mazur, Marcin
Tuzel, Szymon
Snoek, Cees G. M.
Hashemi, Seyyed Hadi
Javed, Omar
Versley, Yannick
Khadivi, Shahram
author_facet Nulli, Matteo
Orshulevich, Vladimir
Bazazo, Tala
Herold, Christian
Kozielski, Michael
Mazur, Marcin
Tuzel, Szymon
Snoek, Cees G. M.
Hashemi, Seyyed Hadi
Javed, Omar
Versley, Yannick
Khadivi, Shahram
contents E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent modelling, yet there is no documented, well-known strategy for adapting them to the attribute-centric, multi-image, and noisy nature of e-commerce data, without sacrificing general performance. In this work, we show through a large-scale experimental study, how targeted adaptation of general VLMs can substantially improve e-commerce performance while preserving broad multimodal capabilities. Furthermore, we propose a novel extensive evaluation suite covering deep product understanding, strict instruction following, and dynamic attribute extraction.
format Preprint
id arxiv_https___arxiv_org_abs_2602_11733
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adapting Vision-Language Models for E-commerce Understanding at Scale
Nulli, Matteo
Orshulevich, Vladimir
Bazazo, Tala
Herold, Christian
Kozielski, Michael
Mazur, Marcin
Tuzel, Szymon
Snoek, Cees G. M.
Hashemi, Seyyed Hadi
Javed, Omar
Versley, Yannick
Khadivi, Shahram
Computer Vision and Pattern Recognition
Artificial Intelligence
E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent modelling, yet there is no documented, well-known strategy for adapting them to the attribute-centric, multi-image, and noisy nature of e-commerce data, without sacrificing general performance. In this work, we show through a large-scale experimental study, how targeted adaptation of general VLMs can substantially improve e-commerce performance while preserving broad multimodal capabilities. Furthermore, we propose a novel extensive evaluation suite covering deep product understanding, strict instruction following, and dynamic attribute extraction.
title Adapting Vision-Language Models for E-commerce Understanding at Scale
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.11733