ERCache: An Efficient and Reliable Caching Framework for Large-Scale User Representations in Meta's Ads System

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Fang, Huang, Yaning, Liang, Dong, Li, Dai, Zhang, Zhongke, Wang, Kai, Xin, Xiao, Aboelela, Abdallah, Jiang, Zheliang, Wang, Yang, Song, Jeff, Zhang, Wei, Liang, Chen, Li, Huayu, Sun, ChongLin, Yang, Hang, Qu, Lei, Shu, Zhan, Yuan, Mindi, Maccherani, Emanuele, Hayat, Taha, Guo, John, Puvvada, Varna, Pashkevich, Uladzimir
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914968178262016
author Zhou, Fang
Huang, Yaning
Liang, Dong
Li, Dai
Zhang, Zhongke
Wang, Kai
Xin, Xiao
Aboelela, Abdallah
Jiang, Zheliang
Wang, Yang
Song, Jeff
Zhang, Wei
Liang, Chen
Li, Huayu
Sun, ChongLin
Yang, Hang
Qu, Lei
Shu, Zhan
Yuan, Mindi
Maccherani, Emanuele
Hayat, Taha
Guo, John
Puvvada, Varna
Pashkevich, Uladzimir
author_facet Zhou, Fang
Huang, Yaning
Liang, Dong
Li, Dai
Zhang, Zhongke
Wang, Kai
Xin, Xiao
Aboelela, Abdallah
Jiang, Zheliang
Wang, Yang
Song, Jeff
Zhang, Wei
Liang, Chen
Li, Huayu
Sun, ChongLin
Yang, Hang
Qu, Lei
Shu, Zhan
Yuan, Mindi
Maccherani, Emanuele
Hayat, Taha
Guo, John
Puvvada, Varna
Pashkevich, Uladzimir
contents The increasing complexity of deep learning models used for calculating user representations presents significant challenges, particularly with limited computational resources and strict service-level agreements (SLAs). Previous research efforts have focused on optimizing model inference but have overlooked a critical question: is it necessary to perform user model inference for every ad request in large-scale social networks? To address this question and these challenges, we first analyze user access patterns at Meta and find that most user model inferences occur within a short timeframe. T his observation reveals a triangular relationship among model complexity, embedding freshness, and service SLAs. Building on this insight, we designed, implemented, and evaluated ERCache, an efficient and robust caching framework for large-scale user representations in ads recommendation systems on social networks. ERCache categorizes cache into direct and failover types and applies customized settings and eviction policies for each model, effectively balancing model complexity, embedding freshness, and service SLAs, even considering the staleness introduced by caching. ERCache has been deployed at Meta for over six months, supporting more than 30 ranking models while efficiently conserving computational resources and complying with service SLA requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06497
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ERCache: An Efficient and Reliable Caching Framework for Large-Scale User Representations in Meta's Ads System
Zhou, Fang
Huang, Yaning
Liang, Dong
Li, Dai
Zhang, Zhongke
Wang, Kai
Xin, Xiao
Aboelela, Abdallah
Jiang, Zheliang
Wang, Yang
Song, Jeff
Zhang, Wei
Liang, Chen
Li, Huayu
Sun, ChongLin
Yang, Hang
Qu, Lei
Shu, Zhan
Yuan, Mindi
Maccherani, Emanuele
Hayat, Taha
Guo, John
Puvvada, Varna
Pashkevich, Uladzimir
Information Retrieval
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
The increasing complexity of deep learning models used for calculating user representations presents significant challenges, particularly with limited computational resources and strict service-level agreements (SLAs). Previous research efforts have focused on optimizing model inference but have overlooked a critical question: is it necessary to perform user model inference for every ad request in large-scale social networks? To address this question and these challenges, we first analyze user access patterns at Meta and find that most user model inferences occur within a short timeframe. T his observation reveals a triangular relationship among model complexity, embedding freshness, and service SLAs. Building on this insight, we designed, implemented, and evaluated ERCache, an efficient and robust caching framework for large-scale user representations in ads recommendation systems on social networks. ERCache categorizes cache into direct and failover types and applies customized settings and eviction policies for each model, effectively balancing model complexity, embedding freshness, and service SLAs, even considering the staleness introduced by caching. ERCache has been deployed at Meta for over six months, supporting more than 30 ranking models while efficiently conserving computational resources and complying with service SLA requirements.
title ERCache: An Efficient and Reliable Caching Framework for Large-Scale User Representations in Meta's Ads System
topic Information Retrieval
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2410.06497