Learning Unified User Quantized Tokenizers for User Representation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Chuan, Chen, Yang, Huang, Wuliang, Zheng, Tianyi, Chen, Jianhu, Dou, Bin, Luo, Yice, Zhu, Yun, Wang, Baokun, Liu, Yongchao, Fu, Xing, Cheng, Yu, Hong, Chuntao, Wang, Weiqiang, Yao, Xin-Wei, Xie, Zhongle
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918151253393408
author He, Chuan
Chen, Yang
Huang, Wuliang
Zheng, Tianyi
Chen, Jianhu
Dou, Bin
Luo, Yice
Zhu, Yun
Wang, Baokun
Liu, Yongchao
Fu, Xing
Cheng, Yu
Hong, Chuntao
Wang, Weiqiang
Yao, Xin-Wei
Xie, Zhongle
author_facet He, Chuan
Chen, Yang
Huang, Wuliang
Zheng, Tianyi
Chen, Jianhu
Dou, Bin
Luo, Yice
Zhu, Yun
Wang, Baokun
Liu, Yongchao
Fu, Xing
Cheng, Yu
Hong, Chuntao
Wang, Weiqiang
Yao, Xin-Wei
Xie, Zhongle
contents Multi-source user representation learning plays a critical role in enabling personalized services on web platforms (e.g., Alipay). While prior works have adopted late-fusion strategies to combine heterogeneous data sources, they suffer from three key limitations: lack of unified representation frameworks, scalability and storage issues in data compression, and inflexible cross-task generalization. To address these challenges, we propose U2QT (Unified User Quantized Tokenizers), a novel framework that integrates cross-domain knowledge transfer with early fusion of heterogeneous domains. Our framework employs a two-stage architecture: first, we use the Qwen3 Embedding model to derive a compact yet expressive feature representation; second, a multi-view RQ-VAE discretizes causal embeddings into compact tokens through shared and source-specific codebooks, enabling efficient storage while maintaining semantic coherence. Experimental results showcase U2QT's advantages across diverse downstream tasks, outperforming task-specific baselines in future behavior prediction and recommendation tasks while achieving efficiency gains in storage and computation. The unified tokenization framework enables seamless integration with language models and supports industrial-scale applications.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00956
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Unified User Quantized Tokenizers for User Representation
He, Chuan
Chen, Yang
Huang, Wuliang
Zheng, Tianyi
Chen, Jianhu
Dou, Bin
Luo, Yice
Zhu, Yun
Wang, Baokun
Liu, Yongchao
Fu, Xing
Cheng, Yu
Hong, Chuntao
Wang, Weiqiang
Yao, Xin-Wei
Xie, Zhongle
Machine Learning
Artificial Intelligence
Information Retrieval
Multi-source user representation learning plays a critical role in enabling personalized services on web platforms (e.g., Alipay). While prior works have adopted late-fusion strategies to combine heterogeneous data sources, they suffer from three key limitations: lack of unified representation frameworks, scalability and storage issues in data compression, and inflexible cross-task generalization. To address these challenges, we propose U2QT (Unified User Quantized Tokenizers), a novel framework that integrates cross-domain knowledge transfer with early fusion of heterogeneous domains. Our framework employs a two-stage architecture: first, we use the Qwen3 Embedding model to derive a compact yet expressive feature representation; second, a multi-view RQ-VAE discretizes causal embeddings into compact tokens through shared and source-specific codebooks, enabling efficient storage while maintaining semantic coherence. Experimental results showcase U2QT's advantages across diverse downstream tasks, outperforming task-specific baselines in future behavior prediction and recommendation tasks while achieving efficiency gains in storage and computation. The unified tokenization framework enables seamless integration with language models and supports industrial-scale applications.
title Learning Unified User Quantized Tokenizers for User Representation
topic Machine Learning
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2508.00956