Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yuhao, Pan, Junwei, Li, Xinhang, Wang, Maolin, Wang, Yuan, Liu, Yue, Liu, Dapeng, Jiang, Jie, Zhao, Xiangyu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911133830479872
author Wang, Yuhao
Pan, Junwei
Li, Xinhang
Wang, Maolin
Wang, Yuan
Liu, Yue
Liu, Dapeng
Jiang, Jie
Zhao, Xiangyu
author_facet Wang, Yuhao
Pan, Junwei
Li, Xinhang
Wang, Maolin
Wang, Yuan
Liu, Yue
Liu, Dapeng
Jiang, Jie
Zhao, Xiangyu
contents Sequential recommendation (SR) aims to capture users' dynamic interests and sequential patterns based on their historical interactions. Recently, the powerful capabilities of large language models (LLMs) have driven their adoption in SR. However, we identify two critical challenges in existing LLM-based SR methods: 1) embedding collapse when incorporating pre-trained collaborative embeddings and 2) catastrophic forgetting of quantized embeddings when utilizing semantic IDs. These issues dampen the model scalability and lead to suboptimal recommendation performance. Therefore, based on LLMs like Llama3-8B-instruct, we introduce a novel SR framework named MME-SID, which integrates multimodal embeddings and quantized embeddings to mitigate embedding collapse. Additionally, we propose a Multimodal Residual Quantized Variational Autoencoder (MM-RQ-VAE) with maximum mean discrepancy as the reconstruction loss and contrastive learning for alignment, which effectively preserve intra-modal distance information and capture inter-modal correlations, respectively. To further alleviate catastrophic forgetting, we initialize the model with the trained multimodal code embeddings. Finally, we fine-tune the LLM efficiently using LoRA in a multimodal frequency-aware fusion manner. Extensive experiments on three public datasets validate the superior performance of MME-SID thanks to its capability to mitigate embedding collapse and catastrophic forgetting. The implementation code and datasets are publicly available for reproduction: https://github.com/Applied-Machine-Learning-Lab/MME-SID.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02017
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs
Wang, Yuhao
Pan, Junwei
Li, Xinhang
Wang, Maolin
Wang, Yuan
Liu, Yue
Liu, Dapeng
Jiang, Jie
Zhao, Xiangyu
Information Retrieval
Artificial Intelligence
Sequential recommendation (SR) aims to capture users' dynamic interests and sequential patterns based on their historical interactions. Recently, the powerful capabilities of large language models (LLMs) have driven their adoption in SR. However, we identify two critical challenges in existing LLM-based SR methods: 1) embedding collapse when incorporating pre-trained collaborative embeddings and 2) catastrophic forgetting of quantized embeddings when utilizing semantic IDs. These issues dampen the model scalability and lead to suboptimal recommendation performance. Therefore, based on LLMs like Llama3-8B-instruct, we introduce a novel SR framework named MME-SID, which integrates multimodal embeddings and quantized embeddings to mitigate embedding collapse. Additionally, we propose a Multimodal Residual Quantized Variational Autoencoder (MM-RQ-VAE) with maximum mean discrepancy as the reconstruction loss and contrastive learning for alignment, which effectively preserve intra-modal distance information and capture inter-modal correlations, respectively. To further alleviate catastrophic forgetting, we initialize the model with the trained multimodal code embeddings. Finally, we fine-tune the LLM efficiently using LoRA in a multimodal frequency-aware fusion manner. Extensive experiments on three public datasets validate the superior performance of MME-SID thanks to its capability to mitigate embedding collapse and catastrophic forgetting. The implementation code and datasets are publicly available for reproduction: https://github.com/Applied-Machine-Learning-Lab/MME-SID.
title Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2509.02017