Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kang, Yue, Huang, Zhuoyi, Schussheim, Benji, Licon, Diana, Atia, Dina, Cao, Shixing, Danovitch, Jacob, Kim, Kunho, Norcilien, Billy, Karpman, Jonah, Sayed, Mahmound, Taylor, Mike, Sun, Tao, Metrikov, Pavel, Agarwal, Vipul, Quirk, Chris, Wang, Ye-Yi, Craswell, Nick, Shaffer, Irene, Chen, Tianwei, Vesal, Sulaiman, Srinivasan, Soundar
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914236778676224
author Kang, Yue
Huang, Zhuoyi
Schussheim, Benji
Licon, Diana
Atia, Dina
Cao, Shixing
Danovitch, Jacob
Kim, Kunho
Norcilien, Billy
Karpman, Jonah
Sayed, Mahmound
Taylor, Mike
Sun, Tao
Metrikov, Pavel
Agarwal, Vipul
Quirk, Chris
Wang, Ye-Yi
Craswell, Nick
Shaffer, Irene
Chen, Tianwei
Vesal, Sulaiman
Srinivasan, Soundar
author_facet Kang, Yue
Huang, Zhuoyi
Schussheim, Benji
Licon, Diana
Atia, Dina
Cao, Shixing
Danovitch, Jacob
Kim, Kunho
Norcilien, Billy
Karpman, Jonah
Sayed, Mahmound
Taylor, Mike
Sun, Tao
Metrikov, Pavel
Agarwal, Vipul
Quirk, Chris
Wang, Ye-Yi
Craswell, Nick
Shaffer, Irene
Chen, Tianwei
Vesal, Sulaiman
Srinivasan, Soundar
contents In enterprise search, building high-quality datasets at scale remains a central challenge due to the difficulty of acquiring labeled data. To resolve this challenge, we propose an efficient approach to fine-tune small language models (SLMs) for accurate relevance labeling, enabling high-throughput, domain-specific labeling comparable or even better in quality to that of state-of-the-art large language models (LLMs). To overcome the lack of high-quality and accessible datasets in the enterprise domain, our method leverages on synthetic data generation. Specifically, we employ an LLM to synthesize realistic enterprise queries from a seed document, apply BM25 to retrieve hard negatives, and use a teacher LLM to assign relevance scores. The resulting dataset is then distilled into an SLM, producing a compact relevance labeler. We evaluate our approach on a high-quality benchmark consisting of 923 enterprise query-document pairs annotated by trained human annotators, and show that the distilled SLM achieves agreement with human judgments on par with or better than the teacher LLM. Furthermore, our fine-tuned labeler substantially improves throughput, achieving 17 times increase while also being 19 times more cost-effective. This approach enables scalable and cost-effective relevance labeling for enterprise-scale retrieval applications, supporting rapid offline evaluation and iteration in real-world settings.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03211
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
Kang, Yue
Huang, Zhuoyi
Schussheim, Benji
Licon, Diana
Atia, Dina
Cao, Shixing
Danovitch, Jacob
Kim, Kunho
Norcilien, Billy
Karpman, Jonah
Sayed, Mahmound
Taylor, Mike
Sun, Tao
Metrikov, Pavel
Agarwal, Vipul
Quirk, Chris
Wang, Ye-Yi
Craswell, Nick
Shaffer, Irene
Chen, Tianwei
Vesal, Sulaiman
Srinivasan, Soundar
Information Retrieval
Artificial Intelligence
Computation and Language
In enterprise search, building high-quality datasets at scale remains a central challenge due to the difficulty of acquiring labeled data. To resolve this challenge, we propose an efficient approach to fine-tune small language models (SLMs) for accurate relevance labeling, enabling high-throughput, domain-specific labeling comparable or even better in quality to that of state-of-the-art large language models (LLMs). To overcome the lack of high-quality and accessible datasets in the enterprise domain, our method leverages on synthetic data generation. Specifically, we employ an LLM to synthesize realistic enterprise queries from a seed document, apply BM25 to retrieve hard negatives, and use a teacher LLM to assign relevance scores. The resulting dataset is then distilled into an SLM, producing a compact relevance labeler. We evaluate our approach on a high-quality benchmark consisting of 923 enterprise query-document pairs annotated by trained human annotators, and show that the distilled SLM achieves agreement with human judgments on par with or better than the teacher LLM. Furthermore, our fine-tuned labeler substantially improves throughput, achieving 17 times increase while also being 19 times more cost-effective. This approach enables scalable and cost-effective relevance labeling for enterprise-scale retrieval applications, supporting rapid offline evaluation and iteration in real-world settings.
title Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.03211