Memento: Personalized RAG-Style Long-Retention Data Scaling for META Ads Recommendation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914593837678592 |
|---|---|
| author | Chen, Xiaoyu Wang, Ruichen Di, Jieming Feng, Suofei Abrar, Nafis Kumari, Lilly Tsui, Tony Liu, Yilin Lu, Yu Patapati, Sowmya Xiong, Junwei Yang, Qiao Sun, Dorothy Cao, Yang Chen, Victor Chen, Pan Sundarkumar, Ramsundar Singh, Shivendra Pratap Overwijk, Arnold Leng, Ling Ramasamy, Dinesh Reddy, Sri Malkin, Robert Pandey, Sandeep |
| author_facet | Chen, Xiaoyu Wang, Ruichen Di, Jieming Feng, Suofei Abrar, Nafis Kumari, Lilly Tsui, Tony Liu, Yilin Lu, Yu Patapati, Sowmya Xiong, Junwei Yang, Qiao Sun, Dorothy Cao, Yang Chen, Victor Chen, Pan Sundarkumar, Ramsundar Singh, Shivendra Pratap Overwijk, Arnold Leng, Ling Ramasamy, Dinesh Reddy, Sri Malkin, Robert Pandey, Sandeep |
| contents | Modeling of long history data suffers from long-context window attention dilution, system efficiency and catastrophic forgetting problems, where naive linear scaling approach like LastN would fail. We introduce Memento, a personalized retrieval-augmented framework that treats historical user engagements as a document corpus and ad requests as queries, retrieving relevant interactions via Maximal Marginal Relevance (MMR) to balance similarity with diversity. We identify two complementary applications: Representation Memento, which retrieves historical embeddings for feature augmentation, and Data Memento, which retrieves past training examples for multipass training. Through infrastructure co-design -- temporal chunking, INT8 quantization, and asynchronous serving -- Memento achieves 5-10$\times$ resource efficiency over linear scaling. Memento processes daily requests with sub-10ms latency, yielding 0.25-0.3% Normalized Entropy gain on both click-through and conversion prediction. In production, Memento delivers a 1% CTR lift on Facebook Feed and Reels and a 1.2% CVR lift, scaling personalization to 365+ days of history. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_24051 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Memento: Personalized RAG-Style Long-Retention Data Scaling for META Ads Recommendation Chen, Xiaoyu Wang, Ruichen Di, Jieming Feng, Suofei Abrar, Nafis Kumari, Lilly Tsui, Tony Liu, Yilin Lu, Yu Patapati, Sowmya Xiong, Junwei Yang, Qiao Sun, Dorothy Cao, Yang Chen, Victor Chen, Pan Sundarkumar, Ramsundar Singh, Shivendra Pratap Overwijk, Arnold Leng, Ling Ramasamy, Dinesh Reddy, Sri Malkin, Robert Pandey, Sandeep Information Retrieval Modeling of long history data suffers from long-context window attention dilution, system efficiency and catastrophic forgetting problems, where naive linear scaling approach like LastN would fail. We introduce Memento, a personalized retrieval-augmented framework that treats historical user engagements as a document corpus and ad requests as queries, retrieving relevant interactions via Maximal Marginal Relevance (MMR) to balance similarity with diversity. We identify two complementary applications: Representation Memento, which retrieves historical embeddings for feature augmentation, and Data Memento, which retrieves past training examples for multipass training. Through infrastructure co-design -- temporal chunking, INT8 quantization, and asynchronous serving -- Memento achieves 5-10$\times$ resource efficiency over linear scaling. Memento processes daily requests with sub-10ms latency, yielding 0.25-0.3% Normalized Entropy gain on both click-through and conversion prediction. In production, Memento delivers a 1% CTR lift on Facebook Feed and Reels and a 1.2% CVR lift, scaling personalization to 365+ days of history. |
| title | Memento: Personalized RAG-Style Long-Retention Data Scaling for META Ads Recommendation |
| topic | Information Retrieval |
| url | https://arxiv.org/abs/2605.24051 |