RASTP: Representation-Aware Semantic Token Pruning for Generative Recommendation with Semantic Identifiers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhan, Tianyu, Fu, Kairui, Lv, Zheqi, Zhang, Shengyu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918304610779136
author Zhan, Tianyu
Fu, Kairui
Lv, Zheqi
Zhang, Shengyu
author_facet Zhan, Tianyu
Fu, Kairui
Lv, Zheqi
Zhang, Shengyu
contents Generative recommendation systems typically leverage Semantic Identifiers (SIDs), which represent each item as a sequence of tokens that encode semantic information. However, representing item ID with multiple SIDs significantly increases input sequence length, which is a major determinant of computational complexity and memory consumption. While existing efforts primarily focus on optimizing attention computation and KV cache, we propose RASTP (Representation-Aware Semantic Token Pruning), which directly prunes less informative tokens in the input sequence. Specifically, RASTP evaluates token importance by combining semantic saliency, measured via representation magnitude, and attention centrality, derived from cumulative attention weights. Since RASTP dynamically prunes low-information or irrelevant semantic tokens, experiments on three real-world Amazon datasets show that RASTP reduces training time by 26.7\%, while maintaining or slightly improving recommendation performance. The code has been open-sourced at https://github.com/Yuzt-zju/RASTP.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16943
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RASTP: Representation-Aware Semantic Token Pruning for Generative Recommendation with Semantic Identifiers
Zhan, Tianyu
Fu, Kairui
Lv, Zheqi
Zhang, Shengyu
Information Retrieval
Artificial Intelligence
Generative recommendation systems typically leverage Semantic Identifiers (SIDs), which represent each item as a sequence of tokens that encode semantic information. However, representing item ID with multiple SIDs significantly increases input sequence length, which is a major determinant of computational complexity and memory consumption. While existing efforts primarily focus on optimizing attention computation and KV cache, we propose RASTP (Representation-Aware Semantic Token Pruning), which directly prunes less informative tokens in the input sequence. Specifically, RASTP evaluates token importance by combining semantic saliency, measured via representation magnitude, and attention centrality, derived from cumulative attention weights. Since RASTP dynamically prunes low-information or irrelevant semantic tokens, experiments on three real-world Amazon datasets show that RASTP reduces training time by 26.7\%, while maintaining or slightly improving recommendation performance. The code has been open-sourced at https://github.com/Yuzt-zju/RASTP.
title RASTP: Representation-Aware Semantic Token Pruning for Generative Recommendation with Semantic Identifiers
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2511.16943