GraphHash: Graph Clustering Enables Parameter Efficiency in Recommender Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Xinyi, Loveland, Donald, Chen, Runjin, Liu, Yozen, Chen, Xin, Neves, Leonardo, Jadbabaie, Ali, Ju, Clark Mingxuan, Shah, Neil, Zhao, Tong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929703270481920
author Wu, Xinyi
Loveland, Donald
Chen, Runjin
Liu, Yozen
Chen, Xin
Neves, Leonardo
Jadbabaie, Ali
Ju, Clark Mingxuan
Shah, Neil
Zhao, Tong
author_facet Wu, Xinyi
Loveland, Donald
Chen, Runjin
Liu, Yozen
Chen, Xin
Neves, Leonardo
Jadbabaie, Ali
Ju, Clark Mingxuan
Shah, Neil
Zhao, Tong
contents Deep recommender systems rely heavily on large embedding tables to handle high-cardinality categorical features such as user/item identifiers, and face significant memory constraints at scale. To tackle this challenge, hashing techniques are often employed to map multiple entities to the same embedding and thus reduce the size of the embedding tables. Concurrently, graph-based collaborative signals have emerged as powerful tools in recommender systems, yet their potential for optimizing embedding table reduction remains unexplored. This paper introduces GraphHash, the first graph-based approach that leverages modularity-based bipartite graph clustering on user-item interaction graphs to reduce embedding table sizes. We demonstrate that the modularity objective has a theoretical connection to message-passing, which provides a foundation for our method. By employing fast clustering algorithms, GraphHash serves as a computationally efficient proxy for message-passing during preprocessing and a plug-and-play graph-based alternative to traditional ID hashing. Extensive experiments show that GraphHash substantially outperforms diverse hashing baselines on both retrieval and click-through-rate prediction tasks. In particular, GraphHash achieves on average a 101.52% improvement in recall when reducing the embedding table size by more than 75%, highlighting the value of graph-based collaborative information for model reduction. Our code is available at https://github.com/snap-research/GraphHash.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17245
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GraphHash: Graph Clustering Enables Parameter Efficiency in Recommender Systems
Wu, Xinyi
Loveland, Donald
Chen, Runjin
Liu, Yozen
Chen, Xin
Neves, Leonardo
Jadbabaie, Ali
Ju, Clark Mingxuan
Shah, Neil
Zhao, Tong
Information Retrieval
Social and Information Networks
Deep recommender systems rely heavily on large embedding tables to handle high-cardinality categorical features such as user/item identifiers, and face significant memory constraints at scale. To tackle this challenge, hashing techniques are often employed to map multiple entities to the same embedding and thus reduce the size of the embedding tables. Concurrently, graph-based collaborative signals have emerged as powerful tools in recommender systems, yet their potential for optimizing embedding table reduction remains unexplored. This paper introduces GraphHash, the first graph-based approach that leverages modularity-based bipartite graph clustering on user-item interaction graphs to reduce embedding table sizes. We demonstrate that the modularity objective has a theoretical connection to message-passing, which provides a foundation for our method. By employing fast clustering algorithms, GraphHash serves as a computationally efficient proxy for message-passing during preprocessing and a plug-and-play graph-based alternative to traditional ID hashing. Extensive experiments show that GraphHash substantially outperforms diverse hashing baselines on both retrieval and click-through-rate prediction tasks. In particular, GraphHash achieves on average a 101.52% improvement in recall when reducing the embedding table size by more than 75%, highlighting the value of graph-based collaborative information for model reduction. Our code is available at https://github.com/snap-research/GraphHash.
title GraphHash: Graph Clustering Enables Parameter Efficiency in Recommender Systems
topic Information Retrieval
Social and Information Networks
url https://arxiv.org/abs/2412.17245