Make It Long, Keep It Fast: End-to-End 10K Long User Behavior Sequence Modeling for Billion-Scale Douyin Recommendation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guan, Lin, Yang, Jia-Qi, Zhao, Zhishan, Zhang, Beichuan, Sun, Bo, Luo, Xuanyuan, Ni, Jinan, Li, Xiaowen, Qi, Yuhang, Fan, Zhifang, Wang, Hangyu, Chen, Qiwei, Cheng, Yi, Zhang, Feng, Yang, Xiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911695374385152
author Guan, Lin
Yang, Jia-Qi
Zhao, Zhishan
Zhang, Beichuan
Sun, Bo
Luo, Xuanyuan
Ni, Jinan
Li, Xiaowen
Qi, Yuhang
Fan, Zhifang
Wang, Hangyu
Chen, Qiwei
Cheng, Yi
Zhang, Feng
Yang, Xiao
author_facet Guan, Lin
Yang, Jia-Qi
Zhao, Zhishan
Zhang, Beichuan
Sun, Bo
Luo, Xuanyuan
Ni, Jinan
Li, Xiaowen
Qi, Yuhang
Fan, Zhifang
Wang, Hangyu
Chen, Qiwei
Cheng, Yi
Zhang, Feng
Yang, Xiao
contents Short-video recommenders such as Douyin must exploit extremely long user behavior histories without breaking latency or cost budgets. We present an end-to-end industrial recommender system that scales long-sequence recommendation modeling to 10K-length histories in production. First, we introduce Stacked Target-to-History Cross Attention (STCA), which replaces history self-attention with stacked cross-attention from the target to the history, reducing complexity from quadratic to linear in sequence length and enabling efficient end-to-end training over long user behavior sequences. Second, we propose Request Level Batching (RLB), a user-centric batching scheme that aggregates multiple targets for the same user/request to share the user-side encoding, substantially lowering sequence-related storage, communication, and compute without changing the learning objective. Third, we design a length-extrapolative training strategy -- train on shorter windows, infer on much longer ones -- so the model generalizes to 10K-scale histories without additional training cost. Across offline and online experiments, we observe predictable, monotonic gains as we scale history length and model capacity, mirroring the scaling law behavior observed in large language models. Deployed at full traffic on Douyin, our system delivers significant improvements on key engagement metrics while meeting production latency, demonstrating a practical path to scaling end-to-end ultra-long sequence recommendation to the 10K regime.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06077
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Make It Long, Keep It Fast: End-to-End 10K Long User Behavior Sequence Modeling for Billion-Scale Douyin Recommendation
Guan, Lin
Yang, Jia-Qi
Zhao, Zhishan
Zhang, Beichuan
Sun, Bo
Luo, Xuanyuan
Ni, Jinan
Li, Xiaowen
Qi, Yuhang
Fan, Zhifang
Wang, Hangyu
Chen, Qiwei
Cheng, Yi
Zhang, Feng
Yang, Xiao
Machine Learning
Information Retrieval
Short-video recommenders such as Douyin must exploit extremely long user behavior histories without breaking latency or cost budgets. We present an end-to-end industrial recommender system that scales long-sequence recommendation modeling to 10K-length histories in production. First, we introduce Stacked Target-to-History Cross Attention (STCA), which replaces history self-attention with stacked cross-attention from the target to the history, reducing complexity from quadratic to linear in sequence length and enabling efficient end-to-end training over long user behavior sequences. Second, we propose Request Level Batching (RLB), a user-centric batching scheme that aggregates multiple targets for the same user/request to share the user-side encoding, substantially lowering sequence-related storage, communication, and compute without changing the learning objective. Third, we design a length-extrapolative training strategy -- train on shorter windows, infer on much longer ones -- so the model generalizes to 10K-scale histories without additional training cost. Across offline and online experiments, we observe predictable, monotonic gains as we scale history length and model capacity, mirroring the scaling law behavior observed in large language models. Deployed at full traffic on Douyin, our system delivers significant improvements on key engagement metrics while meeting production latency, demonstrating a practical path to scaling end-to-end ultra-long sequence recommendation to the 10K regime.
title Make It Long, Keep It Fast: End-to-End 10K Long User Behavior Sequence Modeling for Billion-Scale Douyin Recommendation
topic Machine Learning
Information Retrieval
url https://arxiv.org/abs/2511.06077