Purely Semantic Indexing for LLM-based Generative Recommendation and Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ruohan, Li, Jiacheng, McAuley, Julian, Hou, Yupeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914048628490240
author Zhang, Ruohan
Li, Jiacheng
McAuley, Julian
Hou, Yupeng
author_facet Zhang, Ruohan
Li, Jiacheng
McAuley, Julian
Hou, Yupeng
contents Semantic identifiers (IDs) have proven effective in adapting large language models for generative recommendation and retrieval. However, existing methods often suffer from semantic ID conflicts, where semantically similar documents (or items) are assigned identical IDs. A common strategy to avoid conflicts is to append a non-semantic token to distinguish them, which introduces randomness and expands the search space, therefore hurting performance. In this paper, we propose purely semantic indexing to generate unique, semantic-preserving IDs without appending non-semantic tokens. We enable unique ID assignment by relaxing the strict nearest-centroid selection and introduce two model-agnostic algorithms: exhaustive candidate matching (ECM) and recursive residual searching (RRS). Extensive experiments on sequential recommendation, product search, and document retrieval tasks demonstrate that our methods improve both overall and cold-start performance, highlighting the effectiveness of ensuring ID uniqueness.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16446
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Purely Semantic Indexing for LLM-based Generative Recommendation and Retrieval
Zhang, Ruohan
Li, Jiacheng
McAuley, Julian
Hou, Yupeng
Information Retrieval
Computation and Language
Semantic identifiers (IDs) have proven effective in adapting large language models for generative recommendation and retrieval. However, existing methods often suffer from semantic ID conflicts, where semantically similar documents (or items) are assigned identical IDs. A common strategy to avoid conflicts is to append a non-semantic token to distinguish them, which introduces randomness and expands the search space, therefore hurting performance. In this paper, we propose purely semantic indexing to generate unique, semantic-preserving IDs without appending non-semantic tokens. We enable unique ID assignment by relaxing the strict nearest-centroid selection and introduce two model-agnostic algorithms: exhaustive candidate matching (ECM) and recursive residual searching (RRS). Extensive experiments on sequential recommendation, product search, and document retrieval tasks demonstrate that our methods improve both overall and cold-start performance, highlighting the effectiveness of ensuring ID uniqueness.
title Purely Semantic Indexing for LLM-based Generative Recommendation and Retrieval
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2509.16446