Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shirkavand, Reza, Wei, Xiaokai, Wang, Chen, Hui, Zheng, Huang, Heng, Gong, Michelle
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914465459470336
author Shirkavand, Reza
Wei, Xiaokai
Wang, Chen
Hui, Zheng
Huang, Heng
Gong, Michelle
author_facet Shirkavand, Reza
Wei, Xiaokai
Wang, Chen
Hui, Zheng
Huang, Heng
Gong, Michelle
contents While collaborative filtering delivers predictive accuracy and efficiency, and Large Language Models (LLMs) enable expressive and generalizable reasoning, modern recommendation systems must bring these strengths together. Growing user expectations, such as natural-language queries and transparent explanations, further highlight the need for a unified approach. However, doing so is nontrivial. Collaborative signals are often token-efficient but semantically opaque, while LLMs are semantically rich but struggle to model implicit user preferences when trained only on textual inputs. This paper introduces Item-ID + Oral-language Mixture-of-Experts Language Model (IDIOMoE), which treats item interaction histories as a native dialect within the language space, enabling collaborative signals to be understood in the same way as natural language. By splitting the Feed Forward Network of each block of a pretrained LLM into a separate text expert and an item expert with token-type gating, our method avoids destructive interference between text and catalog modalities. IDIOMoE demonstrates strong recommendation performance across both public and proprietary datasets, while preserving the text understanding of the pretrained model.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05125
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation
Shirkavand, Reza
Wei, Xiaokai
Wang, Chen
Hui, Zheng
Huang, Heng
Gong, Michelle
Computation and Language
Machine Learning
While collaborative filtering delivers predictive accuracy and efficiency, and Large Language Models (LLMs) enable expressive and generalizable reasoning, modern recommendation systems must bring these strengths together. Growing user expectations, such as natural-language queries and transparent explanations, further highlight the need for a unified approach. However, doing so is nontrivial. Collaborative signals are often token-efficient but semantically opaque, while LLMs are semantically rich but struggle to model implicit user preferences when trained only on textual inputs. This paper introduces Item-ID + Oral-language Mixture-of-Experts Language Model (IDIOMoE), which treats item interaction histories as a native dialect within the language space, enabling collaborative signals to be understood in the same way as natural language. By splitting the Feed Forward Network of each block of a pretrained LLM into a separate text expert and an item expert with token-type gating, our method avoids destructive interference between text and catalog modalities. IDIOMoE demonstrates strong recommendation performance across both public and proprietary datasets, while preserving the text understanding of the pretrained model.
title Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.05125