LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Han, Whitworth, Alex, Cheung, Pak Ming, Zhang, Zhenjie, Kamath, Krishna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914149170151424
author Wang, Han
Whitworth, Alex
Cheung, Pak Ming
Zhang, Zhenjie
Kamath, Krishna
author_facet Wang, Han
Whitworth, Alex
Cheung, Pak Ming
Zhang, Zhenjie
Kamath, Krishna
contents Relevance evaluation plays a crucial role in personalized search systems to ensure that search results align with a user's queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present our approach at Pinterest Search to automate relevance evaluation for online experiments using fine-tuned LLMs. We rigorously validate the alignment between LLM-generated judgments and human annotations, demonstrating that LLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging LLM-based labeling further unlocks the opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effect (MDE) in online experiment measurements.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03764
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest
Wang, Han
Whitworth, Alex
Cheung, Pak Ming
Zhang, Zhenjie
Kamath, Krishna
Information Retrieval
Machine Learning
Relevance evaluation plays a crucial role in personalized search systems to ensure that search results align with a user's queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present our approach at Pinterest Search to automate relevance evaluation for online experiments using fine-tuned LLMs. We rigorously validate the alignment between LLM-generated judgments and human annotations, demonstrating that LLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging LLM-based labeling further unlocks the opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effect (MDE) in online experiment measurements.
title LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2509.03764