LiNR: Model Based Neural Retrieval on GPUs at LinkedIn
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917743277637632 |
|---|---|
| author | Borisyuk, Fedor Song, Qingquan Zhou, Mingzhou Parameswaran, Ganesh Arun, Madhu Popuri, Siva Bingol, Tugrul Pei, Zhuotao Lee, Kuang-Hsuan Zheng, Lu Shao, Qizhan Naqvi, Ali Zhou, Sen Gupta, Aman |
| author_facet | Borisyuk, Fedor Song, Qingquan Zhou, Mingzhou Parameswaran, Ganesh Arun, Madhu Popuri, Siva Bingol, Tugrul Pei, Zhuotao Lee, Kuang-Hsuan Zheng, Lu Shao, Qizhan Naqvi, Ali Zhou, Sen Gupta, Aman |
| contents | This paper introduces LiNR, LinkedIn's large-scale, GPU-based retrieval system. LiNR supports a billion-sized index on GPU models. We discuss our experiences and challenges in creating scalable, differentiable search indexes using TensorFlow and PyTorch at production scale. In LiNR, both items and model weights are integrated into the model binary. Viewing index construction as a form of model training, we describe scaling our system for large indexes, incorporating full scans and efficient filtering. A key focus is on enabling attribute-based pre-filtering for exhaustive GPU searches, addressing the common challenge of post-filtering in KNN searches that often reduces system quality. We further provide multi-embedding retrieval algorithms and strategies for tackling cold start issues in retrieval. Our advancements in supporting larger indexes through quantization are also discussed. We believe LiNR represents one of the industry's first Live-updated model-based retrieval indexes. Applied to out-of-network post recommendations on LinkedIn Feed, LiNR has contributed to a 3% relative increase in professional daily active users. We envisage LiNR as a step towards integrating retrieval and ranking into a single GPU model, simplifying complex infrastructures and enabling end-to-end optimization of the entire differentiable infrastructure through gradient descent. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_13218 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LiNR: Model Based Neural Retrieval on GPUs at LinkedIn Borisyuk, Fedor Song, Qingquan Zhou, Mingzhou Parameswaran, Ganesh Arun, Madhu Popuri, Siva Bingol, Tugrul Pei, Zhuotao Lee, Kuang-Hsuan Zheng, Lu Shao, Qizhan Naqvi, Ali Zhou, Sen Gupta, Aman Machine Learning Artificial Intelligence This paper introduces LiNR, LinkedIn's large-scale, GPU-based retrieval system. LiNR supports a billion-sized index on GPU models. We discuss our experiences and challenges in creating scalable, differentiable search indexes using TensorFlow and PyTorch at production scale. In LiNR, both items and model weights are integrated into the model binary. Viewing index construction as a form of model training, we describe scaling our system for large indexes, incorporating full scans and efficient filtering. A key focus is on enabling attribute-based pre-filtering for exhaustive GPU searches, addressing the common challenge of post-filtering in KNN searches that often reduces system quality. We further provide multi-embedding retrieval algorithms and strategies for tackling cold start issues in retrieval. Our advancements in supporting larger indexes through quantization are also discussed. We believe LiNR represents one of the industry's first Live-updated model-based retrieval indexes. Applied to out-of-network post recommendations on LinkedIn Feed, LiNR has contributed to a 3% relative increase in professional daily active users. We envisage LiNR as a step towards integrating retrieval and ranking into a single GPU model, simplifying complex infrastructures and enabling end-to-end optimization of the entire differentiable infrastructure through gradient descent. |
| title | LiNR: Model Based Neural Retrieval on GPUs at LinkedIn |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2407.13218 |