Efficient Item ID Generation for Large-Scale LLM-based Recommendation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Subbiah, Anushya, Aggarwal, Vikram, Pine, James, Rendle, Steffen, Sayana, Krishna, Su, Kun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909769323773952
author Subbiah, Anushya
Aggarwal, Vikram
Pine, James
Rendle, Steffen
Sayana, Krishna
Su, Kun
author_facet Subbiah, Anushya
Aggarwal, Vikram
Pine, James
Rendle, Steffen
Sayana, Krishna
Su, Kun
contents Integrating product catalogs and user behavior into LLMs can enhance recommendations with broad world knowledge, but the scale of real-world item catalogs, often containing millions of discrete item identifiers (Item IDs), poses a significant challenge. This contrasts with the smaller, tokenized text vocabularies typically used in LLMs. The predominant view within the LLM-based recommendation literature is that it is infeasible to treat item ids as a first class citizen in the LLM and instead some sort of tokenization of an item into multiple tokens is required. However, this creates a key practical bottleneck in serving these models for real-time low-latency applications. Our paper challenges this predominant practice and integrates item ids as first class citizens into the LLM. We provide simple, yet highly effective, novel training and inference modifications that enable single-token representations of items and single-step decoding. Our method shows improvements in recommendation quality (Recall and NDCG) over existing techniques on the Amazon shopping datasets while significantly improving inference efficiency by 5x-14x. Our work offers an efficiency perspective distinct from that of other popular approaches within LLM-based recommendation, potentially inspiring further research and opening up a new direction for integrating IDs into LLMs. Our code is available here https://drive.google.com/file/d/1cUMj37rV0Z1bCWMdhQ6i4q4eTRQLURtC
format Preprint
id arxiv_https___arxiv_org_abs_2509_03746
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Item ID Generation for Large-Scale LLM-based Recommendation
Subbiah, Anushya
Aggarwal, Vikram
Pine, James
Rendle, Steffen
Sayana, Krishna
Su, Kun
Information Retrieval
Integrating product catalogs and user behavior into LLMs can enhance recommendations with broad world knowledge, but the scale of real-world item catalogs, often containing millions of discrete item identifiers (Item IDs), poses a significant challenge. This contrasts with the smaller, tokenized text vocabularies typically used in LLMs. The predominant view within the LLM-based recommendation literature is that it is infeasible to treat item ids as a first class citizen in the LLM and instead some sort of tokenization of an item into multiple tokens is required. However, this creates a key practical bottleneck in serving these models for real-time low-latency applications. Our paper challenges this predominant practice and integrates item ids as first class citizens into the LLM. We provide simple, yet highly effective, novel training and inference modifications that enable single-token representations of items and single-step decoding. Our method shows improvements in recommendation quality (Recall and NDCG) over existing techniques on the Amazon shopping datasets while significantly improving inference efficiency by 5x-14x. Our work offers an efficiency perspective distinct from that of other popular approaches within LLM-based recommendation, potentially inspiring further research and opening up a new direction for integrating IDs into LLMs. Our code is available here https://drive.google.com/file/d/1cUMj37rV0Z1bCWMdhQ6i4q4eTRQLURtC
title Efficient Item ID Generation for Large-Scale LLM-based Recommendation
topic Information Retrieval
url https://arxiv.org/abs/2509.03746