Learning Unified User Quantized Tokenizers for User Representation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918151253393408 |
|---|---|
| author | He, Chuan Chen, Yang Huang, Wuliang Zheng, Tianyi Chen, Jianhu Dou, Bin Luo, Yice Zhu, Yun Wang, Baokun Liu, Yongchao Fu, Xing Cheng, Yu Hong, Chuntao Wang, Weiqiang Yao, Xin-Wei Xie, Zhongle |
| author_facet | He, Chuan Chen, Yang Huang, Wuliang Zheng, Tianyi Chen, Jianhu Dou, Bin Luo, Yice Zhu, Yun Wang, Baokun Liu, Yongchao Fu, Xing Cheng, Yu Hong, Chuntao Wang, Weiqiang Yao, Xin-Wei Xie, Zhongle |
| contents | Multi-source user representation learning plays a critical role in enabling personalized services on web platforms (e.g., Alipay). While prior works have adopted late-fusion strategies to combine heterogeneous data sources, they suffer from three key limitations: lack of unified representation frameworks, scalability and storage issues in data compression, and inflexible cross-task generalization. To address these challenges, we propose U2QT (Unified User Quantized Tokenizers), a novel framework that integrates cross-domain knowledge transfer with early fusion of heterogeneous domains. Our framework employs a two-stage architecture: first, we use the Qwen3 Embedding model to derive a compact yet expressive feature representation; second, a multi-view RQ-VAE discretizes causal embeddings into compact tokens through shared and source-specific codebooks, enabling efficient storage while maintaining semantic coherence. Experimental results showcase U2QT's advantages across diverse downstream tasks, outperforming task-specific baselines in future behavior prediction and recommendation tasks while achieving efficiency gains in storage and computation. The unified tokenization framework enables seamless integration with language models and supports industrial-scale applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_00956 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Learning Unified User Quantized Tokenizers for User Representation He, Chuan Chen, Yang Huang, Wuliang Zheng, Tianyi Chen, Jianhu Dou, Bin Luo, Yice Zhu, Yun Wang, Baokun Liu, Yongchao Fu, Xing Cheng, Yu Hong, Chuntao Wang, Weiqiang Yao, Xin-Wei Xie, Zhongle Machine Learning Artificial Intelligence Information Retrieval Multi-source user representation learning plays a critical role in enabling personalized services on web platforms (e.g., Alipay). While prior works have adopted late-fusion strategies to combine heterogeneous data sources, they suffer from three key limitations: lack of unified representation frameworks, scalability and storage issues in data compression, and inflexible cross-task generalization. To address these challenges, we propose U2QT (Unified User Quantized Tokenizers), a novel framework that integrates cross-domain knowledge transfer with early fusion of heterogeneous domains. Our framework employs a two-stage architecture: first, we use the Qwen3 Embedding model to derive a compact yet expressive feature representation; second, a multi-view RQ-VAE discretizes causal embeddings into compact tokens through shared and source-specific codebooks, enabling efficient storage while maintaining semantic coherence. Experimental results showcase U2QT's advantages across diverse downstream tasks, outperforming task-specific baselines in future behavior prediction and recommendation tasks while achieving efficiency gains in storage and computation. The unified tokenization framework enables seamless integration with language models and supports industrial-scale applications. |
| title | Learning Unified User Quantized Tokenizers for User Representation |
| topic | Machine Learning Artificial Intelligence Information Retrieval |
| url | https://arxiv.org/abs/2508.00956 |