Adapting Vision-Language Models for E-commerce Understanding at Scale
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914324762591232 |
|---|---|
| author | Nulli, Matteo Orshulevich, Vladimir Bazazo, Tala Herold, Christian Kozielski, Michael Mazur, Marcin Tuzel, Szymon Snoek, Cees G. M. Hashemi, Seyyed Hadi Javed, Omar Versley, Yannick Khadivi, Shahram |
| author_facet | Nulli, Matteo Orshulevich, Vladimir Bazazo, Tala Herold, Christian Kozielski, Michael Mazur, Marcin Tuzel, Szymon Snoek, Cees G. M. Hashemi, Seyyed Hadi Javed, Omar Versley, Yannick Khadivi, Shahram |
| contents | E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent modelling, yet there is no documented, well-known strategy for adapting them to the attribute-centric, multi-image, and noisy nature of e-commerce data, without sacrificing general performance. In this work, we show through a large-scale experimental study, how targeted adaptation of general VLMs can substantially improve e-commerce performance while preserving broad multimodal capabilities. Furthermore, we propose a novel extensive evaluation suite covering deep product understanding, strict instruction following, and dynamic attribute extraction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_11733 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Adapting Vision-Language Models for E-commerce Understanding at Scale Nulli, Matteo Orshulevich, Vladimir Bazazo, Tala Herold, Christian Kozielski, Michael Mazur, Marcin Tuzel, Szymon Snoek, Cees G. M. Hashemi, Seyyed Hadi Javed, Omar Versley, Yannick Khadivi, Shahram Computer Vision and Pattern Recognition Artificial Intelligence E-commerce product understanding demands by nature, strong multimodal comprehension from text, images, and structured attributes. General-purpose Vision-Language Models (VLMs) enable generalizable multimodal latent modelling, yet there is no documented, well-known strategy for adapting them to the attribute-centric, multi-image, and noisy nature of e-commerce data, without sacrificing general performance. In this work, we show through a large-scale experimental study, how targeted adaptation of general VLMs can substantially improve e-commerce performance while preserving broad multimodal capabilities. Furthermore, we propose a novel extensive evaluation suite covering deep product understanding, strict instruction following, and dynamic attribute extraction. |
| title | Adapting Vision-Language Models for E-commerce Understanding at Scale |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2602.11733 |