Dataset Condensation for Recommendation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Jiahao, Fan, Wenqi, Chen, Jingfan, Liu, Shengcai, Liu, Qijiong, He, Rui, Li, Qing, Tang, Ke
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909571873767424
author Wu, Jiahao
Fan, Wenqi
Chen, Jingfan
Liu, Shengcai
Liu, Qijiong
He, Rui
Li, Qing
Tang, Ke
author_facet Wu, Jiahao
Fan, Wenqi
Chen, Jingfan
Liu, Shengcai
Liu, Qijiong
He, Rui
Li, Qing
Tang, Ke
contents Training recommendation models on large datasets requires significant time and resources. It is desired to construct concise yet informative datasets for efficient training. Recent advances in dataset condensation show promise in addressing this problem by synthesizing small datasets. However, applying existing methods of dataset condensation to recommendation has limitations: (1) they fail to generate discrete user-item interactions, and (2) they could not preserve users' potential preferences. To address the limitations, we propose a lightweight condensation framework tailored for recommendation (DConRec), focusing on condensing user-item historical interaction sets. Specifically, we model the discrete user-item interactions via a probabilistic approach and design a pre-augmentation module to incorporate the potential preferences of users into the condensed datasets. While the substantial size of datasets leads to costly optimization, we propose a lightweight policy gradient estimation to accelerate the data synthesis. Experimental results on multiple real-world datasets have demonstrated the effectiveness and efficiency of our framework. Besides, we provide a theoretical analysis of the provable convergence of DConRec. Our implementation is available at: https://github.com/JiahaoWuGit/DConRec.
format Preprint
id arxiv_https___arxiv_org_abs_2310_01038
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Dataset Condensation for Recommendation
Wu, Jiahao
Fan, Wenqi
Chen, Jingfan
Liu, Shengcai
Liu, Qijiong
He, Rui
Li, Qing
Tang, Ke
Information Retrieval
Training recommendation models on large datasets requires significant time and resources. It is desired to construct concise yet informative datasets for efficient training. Recent advances in dataset condensation show promise in addressing this problem by synthesizing small datasets. However, applying existing methods of dataset condensation to recommendation has limitations: (1) they fail to generate discrete user-item interactions, and (2) they could not preserve users' potential preferences. To address the limitations, we propose a lightweight condensation framework tailored for recommendation (DConRec), focusing on condensing user-item historical interaction sets. Specifically, we model the discrete user-item interactions via a probabilistic approach and design a pre-augmentation module to incorporate the potential preferences of users into the condensed datasets. While the substantial size of datasets leads to costly optimization, we propose a lightweight policy gradient estimation to accelerate the data synthesis. Experimental results on multiple real-world datasets have demonstrated the effectiveness and efficiency of our framework. Besides, we provide a theoretical analysis of the provable convergence of DConRec. Our implementation is available at: https://github.com/JiahaoWuGit/DConRec.
title Dataset Condensation for Recommendation
topic Information Retrieval
url https://arxiv.org/abs/2310.01038