Pre-training Generative Recommender with Multi-Identifier Item Tokenization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Bowen, Liu, Enze, Chen, Zhongfu, Ma, Zhongrui, Wang, Yue, Zhao, Wayne Xin, Wen, Ji-Rong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912391374045184
author Zheng, Bowen
Liu, Enze
Chen, Zhongfu
Ma, Zhongrui
Wang, Yue
Zhao, Wayne Xin
Wen, Ji-Rong
author_facet Zheng, Bowen
Liu, Enze
Chen, Zhongfu
Ma, Zhongrui
Wang, Yue
Zhao, Wayne Xin
Wen, Ji-Rong
contents Generative recommendation autoregressively generates item identifiers to recommend potential items. Existing methods typically adopt a one-to-one mapping strategy, where each item is represented by a single identifier. However, this scheme poses issues, such as suboptimal semantic modeling for low-frequency items and limited diversity in token sequence data. To overcome these limitations, we propose MTGRec, which leverages Multi-identifier item Tokenization to augment token sequence data for Generative Recommender pre-training. Our approach involves two key innovations: multi-identifier item tokenization and curriculum recommender pre-training. For multi-identifier item tokenization, we leverage the RQ-VAE as the tokenizer backbone and treat model checkpoints from adjacent training epochs as semantically relevant tokenizers. This allows each item to be associated with multiple identifiers, enabling a single user interaction sequence to be converted into several token sequences as different data groups. For curriculum recommender pre-training, we introduce a curriculum learning scheme guided by data influence estimation, dynamically adjusting the sampling probability of each data group during recommender pre-training. After pre-training, we fine-tune the model using a single tokenizer to ensure accurate item identification for recommendation. Extensive experiments on three public benchmark datasets demonstrate that MTGRec significantly outperforms both traditional and generative recommendation baselines in terms of effectiveness and scalability.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04400
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pre-training Generative Recommender with Multi-Identifier Item Tokenization
Zheng, Bowen
Liu, Enze
Chen, Zhongfu
Ma, Zhongrui
Wang, Yue
Zhao, Wayne Xin
Wen, Ji-Rong
Information Retrieval
Artificial Intelligence
Generative recommendation autoregressively generates item identifiers to recommend potential items. Existing methods typically adopt a one-to-one mapping strategy, where each item is represented by a single identifier. However, this scheme poses issues, such as suboptimal semantic modeling for low-frequency items and limited diversity in token sequence data. To overcome these limitations, we propose MTGRec, which leverages Multi-identifier item Tokenization to augment token sequence data for Generative Recommender pre-training. Our approach involves two key innovations: multi-identifier item tokenization and curriculum recommender pre-training. For multi-identifier item tokenization, we leverage the RQ-VAE as the tokenizer backbone and treat model checkpoints from adjacent training epochs as semantically relevant tokenizers. This allows each item to be associated with multiple identifiers, enabling a single user interaction sequence to be converted into several token sequences as different data groups. For curriculum recommender pre-training, we introduce a curriculum learning scheme guided by data influence estimation, dynamically adjusting the sampling probability of each data group during recommender pre-training. After pre-training, we fine-tune the model using a single tokenizer to ensure accurate item identification for recommendation. Extensive experiments on three public benchmark datasets demonstrate that MTGRec significantly outperforms both traditional and generative recommendation baselines in terms of effectiveness and scalability.
title Pre-training Generative Recommender with Multi-Identifier Item Tokenization
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2504.04400