Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhimin, Zhao, Chenyu, Mo, Ka Chun, Jiang, Yunjiang, Lee, Jane H., Mahajan, Khushhall Chandra, Jiang, Ning, Ren, Kai, Li, Jinhui, Yang, Wen-Yun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910074610384896
author Chen, Zhimin
Zhao, Chenyu
Mo, Ka Chun
Jiang, Yunjiang
Lee, Jane H.
Mahajan, Khushhall Chandra
Jiang, Ning
Ren, Kai
Li, Jinhui
Yang, Wen-Yun
author_facet Chen, Zhimin
Zhao, Chenyu
Mo, Ka Chun
Jiang, Yunjiang
Lee, Jane H.
Mahajan, Khushhall Chandra
Jiang, Ning
Ren, Kai
Li, Jinhui
Yang, Wen-Yun
contents Modern large-scale recommendation systems rely heavily on user interaction history sequences to enhance the model performance. The advent of large language models and sequential modeling techniques, particularly transformer-like architectures, has led to significant advancements recently (e.g., HSTU, SIM, and TWIN models). While scaling to ultra-long user histories (10k to 100k items) generally improves model performance, it also creates significant challenges on latency, queries per second (QPS) and GPU cost in industry-scale recommendation systems. Existing models do not adequately address these industrial scalability issues. In this paper, we propose a novel two-stage modeling framework, namely VIrtual Sequential Target Attention (VISTA), which decomposes traditional target attention from a candidate item to user history items into two distinct stages: (1) user history summarization into a few hundred tokens; followed by (2) candidate item attention to those tokens. These summarization token embeddings are then cached in storage system and then utilized as sequence features for downstream model training and inference. This novel design for scalability enables VISTA to scale to lifelong user histories (up to one million items) while keeping downstream training and inference costs fixed, which is essential in industry. Our approach achieves significant improvements in offline and online metrics and has been successfully deployed on an industry leading recommendation platform serving billions of users.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders
Chen, Zhimin
Zhao, Chenyu
Mo, Ka Chun
Jiang, Yunjiang
Lee, Jane H.
Mahajan, Khushhall Chandra
Jiang, Ning
Ren, Kai
Li, Jinhui
Yang, Wen-Yun
Information Retrieval
Machine Learning
Modern large-scale recommendation systems rely heavily on user interaction history sequences to enhance the model performance. The advent of large language models and sequential modeling techniques, particularly transformer-like architectures, has led to significant advancements recently (e.g., HSTU, SIM, and TWIN models). While scaling to ultra-long user histories (10k to 100k items) generally improves model performance, it also creates significant challenges on latency, queries per second (QPS) and GPU cost in industry-scale recommendation systems. Existing models do not adequately address these industrial scalability issues. In this paper, we propose a novel two-stage modeling framework, namely VIrtual Sequential Target Attention (VISTA), which decomposes traditional target attention from a candidate item to user history items into two distinct stages: (1) user history summarization into a few hundred tokens; followed by (2) candidate item attention to those tokens. These summarization token embeddings are then cached in storage system and then utilized as sequence features for downstream model training and inference. This novel design for scalability enables VISTA to scale to lifelong user histories (up to one million items) while keeping downstream training and inference costs fixed, which is essential in industry. Our approach achieves significant improvements in offline and online metrics and has been successfully deployed on an industry leading recommendation platform serving billions of users.
title Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2510.22049