Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rusli, Andre, Ishimoto, Shoma, Akiyama, Sho, Singh, Aman Kumar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912526918221824
author Rusli, Andre
Ishimoto, Shoma
Akiyama, Sho
Singh, Aman Kumar
author_facet Rusli, Andre
Ishimoto, Shoma
Akiyama, Sho
Singh, Aman Kumar
contents Visual search offers an intuitive way for customers to explore diverse product catalogs, particularly in consumer-to-consumer (C2C) marketplaces where listings are often unstructured and visually driven. This paper presents a scalable visual search system deployed in Mercari's C2C marketplace, where end-users act as buyers and sellers. We evaluate recent vision-language models for zero-shot image retrieval and compare their performance with an existing fine-tuned baseline. The system integrates real-time inference and background indexing workflows, supported by a unified embedding pipeline optimized through dimensionality reduction. Offline evaluation using user interaction logs shows that the multilingual SigLIP model outperforms other models across multiple retrieval metrics, achieving a 13.3% increase in nDCG@5 over the baseline. A one-week online A/B test in production further confirms real-world impact, with the treatment group showing substantial gains in engagement and conversion, up to a 40.9% increase in transaction rate via image search. Our findings highlight that recent zero-shot models can serve as a strong and practical baseline for production use, which enables teams to deploy effective visual search systems with minimal overhead, while retaining the flexibility to fine-tune based on future data or domain-specific needs.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace
Rusli, Andre
Ishimoto, Shoma
Akiyama, Sho
Singh, Aman Kumar
Information Retrieval
Artificial Intelligence
Visual search offers an intuitive way for customers to explore diverse product catalogs, particularly in consumer-to-consumer (C2C) marketplaces where listings are often unstructured and visually driven. This paper presents a scalable visual search system deployed in Mercari's C2C marketplace, where end-users act as buyers and sellers. We evaluate recent vision-language models for zero-shot image retrieval and compare their performance with an existing fine-tuned baseline. The system integrates real-time inference and background indexing workflows, supported by a unified embedding pipeline optimized through dimensionality reduction. Offline evaluation using user interaction logs shows that the multilingual SigLIP model outperforms other models across multiple retrieval metrics, achieving a 13.3% increase in nDCG@5 over the baseline. A one-week online A/B test in production further confirms real-world impact, with the treatment group showing substantial gains in engagement and conversion, up to a 40.9% increase in transaction rate via image search. Our findings highlight that recent zero-shot models can serve as a strong and practical baseline for production use, which enables teams to deploy effective visual search systems with minimal overhead, while retaining the flexibility to fine-tune based on future data or domain-specific needs.
title Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2508.05661