ERCache: An Efficient and Reliable Caching Framework for Large-Scale User Representations in Meta's Ads System
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914968178262016 |
|---|---|
| author | Zhou, Fang Huang, Yaning Liang, Dong Li, Dai Zhang, Zhongke Wang, Kai Xin, Xiao Aboelela, Abdallah Jiang, Zheliang Wang, Yang Song, Jeff Zhang, Wei Liang, Chen Li, Huayu Sun, ChongLin Yang, Hang Qu, Lei Shu, Zhan Yuan, Mindi Maccherani, Emanuele Hayat, Taha Guo, John Puvvada, Varna Pashkevich, Uladzimir |
| author_facet | Zhou, Fang Huang, Yaning Liang, Dong Li, Dai Zhang, Zhongke Wang, Kai Xin, Xiao Aboelela, Abdallah Jiang, Zheliang Wang, Yang Song, Jeff Zhang, Wei Liang, Chen Li, Huayu Sun, ChongLin Yang, Hang Qu, Lei Shu, Zhan Yuan, Mindi Maccherani, Emanuele Hayat, Taha Guo, John Puvvada, Varna Pashkevich, Uladzimir |
| contents | The increasing complexity of deep learning models used for calculating user representations presents significant challenges, particularly with limited computational resources and strict service-level agreements (SLAs). Previous research efforts have focused on optimizing model inference but have overlooked a critical question: is it necessary to perform user model inference for every ad request in large-scale social networks? To address this question and these challenges, we first analyze user access patterns at Meta and find that most user model inferences occur within a short timeframe. T his observation reveals a triangular relationship among model complexity, embedding freshness, and service SLAs. Building on this insight, we designed, implemented, and evaluated ERCache, an efficient and robust caching framework for large-scale user representations in ads recommendation systems on social networks. ERCache categorizes cache into direct and failover types and applies customized settings and eviction policies for each model, effectively balancing model complexity, embedding freshness, and service SLAs, even considering the staleness introduced by caching. ERCache has been deployed at Meta for over six months, supporting more than 30 ranking models while efficiently conserving computational resources and complying with service SLA requirements. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_06497 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ERCache: An Efficient and Reliable Caching Framework for Large-Scale User Representations in Meta's Ads System Zhou, Fang Huang, Yaning Liang, Dong Li, Dai Zhang, Zhongke Wang, Kai Xin, Xiao Aboelela, Abdallah Jiang, Zheliang Wang, Yang Song, Jeff Zhang, Wei Liang, Chen Li, Huayu Sun, ChongLin Yang, Hang Qu, Lei Shu, Zhan Yuan, Mindi Maccherani, Emanuele Hayat, Taha Guo, John Puvvada, Varna Pashkevich, Uladzimir Information Retrieval Artificial Intelligence Distributed, Parallel, and Cluster Computing Machine Learning The increasing complexity of deep learning models used for calculating user representations presents significant challenges, particularly with limited computational resources and strict service-level agreements (SLAs). Previous research efforts have focused on optimizing model inference but have overlooked a critical question: is it necessary to perform user model inference for every ad request in large-scale social networks? To address this question and these challenges, we first analyze user access patterns at Meta and find that most user model inferences occur within a short timeframe. T his observation reveals a triangular relationship among model complexity, embedding freshness, and service SLAs. Building on this insight, we designed, implemented, and evaluated ERCache, an efficient and robust caching framework for large-scale user representations in ads recommendation systems on social networks. ERCache categorizes cache into direct and failover types and applies customized settings and eviction policies for each model, effectively balancing model complexity, embedding freshness, and service SLAs, even considering the staleness introduced by caching. ERCache has been deployed at Meta for over six months, supporting more than 30 ranking models while efficiently conserving computational resources and complying with service SLA requirements. |
| title | ERCache: An Efficient and Reliable Caching Framework for Large-Scale User Representations in Meta's Ads System |
| topic | Information Retrieval Artificial Intelligence Distributed, Parallel, and Cluster Computing Machine Learning |
| url | https://arxiv.org/abs/2410.06497 |