Versioned Late Materialization for Ultra-Long Sequence Training in Recommendation Systems at Scale

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Guo, Liang, Song, Ge, Deng, Litao, Sun, Jianhui, Hu, Chufeng, Zhang, Lu, Ma, Zhen, Chen, Shouwei, Liu, Weiran, Sreeshylan, Sarang Masti, Meng, Xiaoxuan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917442235662336
author Guo, Liang
Song, Ge
Deng, Litao
Sun, Jianhui
Hu, Chufeng
Zhang, Lu
Ma, Zhen
Chen, Shouwei
Liu, Weiran
Sreeshylan, Sarang Masti
Meng, Xiaoxuan
author_facet Guo, Liang
Song, Ge
Deng, Litao
Sun, Jianhui
Hu, Chufeng
Zhang, Lu
Ma, Zhen
Chen, Shouwei
Liu, Weiran
Sreeshylan, Sarang Masti
Meng, Xiaoxuan
contents Modern Deep Learning Recommendation Models (DLRMs) follow scaling laws with sequence length, driving the frontier toward ultra-long User Interaction History (UIH). However, the industry-standard "Fat Row" paradigm, which pre-materializes these sequences into every training example, creates a storage and I/O wall where data infrastructure usage exceeds GPU training capacity due to data redundancy that is amplified in multi-tenant environments where models with vastly different sequence length requirements share a union dataset. We present a \emph{versioned late materialization} paradigm that eliminates this redundancy by storing UIH once in a normalized, immutable tier and reconstructing sequences just-in-time during training via lightweight versioned pointers. The system ensures Online-to-Offline (O2O) consistency through a bifurcated protocol that prevents future leakage across both streaming and batch training, while a read-optimized immutable storage layer provides multi-dimensional projection pushdown for heterogeneous model tenants. Disaggregated data preprocessing with pipelined I/O prefetching and data-affinity optimizations masks the latency of training-time sequence reconstruction, keeping training throughput compute-bound by GPUs. Deployed on production DLRMs, the system reduces training data infrastructure resource usage while enabling aggressive sequence length scaling that delivers significant model quality gains, serving as the foundational data infrastructure for modern recommendation model architectures, including HSTU and ULTRA-HSTU.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24806
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Versioned Late Materialization for Ultra-Long Sequence Training in Recommendation Systems at Scale
Guo, Liang
Song, Ge
Deng, Litao
Sun, Jianhui
Hu, Chufeng
Zhang, Lu
Ma, Zhen
Chen, Shouwei
Liu, Weiran
Sreeshylan, Sarang Masti
Meng, Xiaoxuan
Information Retrieval
Artificial Intelligence
Databases
Modern Deep Learning Recommendation Models (DLRMs) follow scaling laws with sequence length, driving the frontier toward ultra-long User Interaction History (UIH). However, the industry-standard "Fat Row" paradigm, which pre-materializes these sequences into every training example, creates a storage and I/O wall where data infrastructure usage exceeds GPU training capacity due to data redundancy that is amplified in multi-tenant environments where models with vastly different sequence length requirements share a union dataset. We present a \emph{versioned late materialization} paradigm that eliminates this redundancy by storing UIH once in a normalized, immutable tier and reconstructing sequences just-in-time during training via lightweight versioned pointers. The system ensures Online-to-Offline (O2O) consistency through a bifurcated protocol that prevents future leakage across both streaming and batch training, while a read-optimized immutable storage layer provides multi-dimensional projection pushdown for heterogeneous model tenants. Disaggregated data preprocessing with pipelined I/O prefetching and data-affinity optimizations masks the latency of training-time sequence reconstruction, keeping training throughput compute-bound by GPUs. Deployed on production DLRMs, the system reduces training data infrastructure resource usage while enabling aggressive sequence length scaling that delivers significant model quality gains, serving as the foundational data infrastructure for modern recommendation model architectures, including HSTU and ULTRA-HSTU.
title Versioned Late Materialization for Ultra-Long Sequence Training in Recommendation Systems at Scale
topic Information Retrieval
Artificial Intelligence
Databases
url https://arxiv.org/abs/2604.24806