Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yifan, Wen, Qianfeng, Liang, Jiazhou, Zhao, Mark, Cui, Justin, Korikov, Anton, Toroghi, Armin, Kim, Junyoung, Sanner, Scott
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909880745459712
author Liu, Yifan
Wen, Qianfeng
Liang, Jiazhou
Zhao, Mark
Cui, Justin
Korikov, Anton
Toroghi, Armin
Kim, Junyoung
Sanner, Scott
author_facet Liu, Yifan
Wen, Qianfeng
Liang, Jiazhou
Zhao, Mark
Cui, Justin
Korikov, Anton
Toroghi, Armin
Kim, Junyoung
Sanner, Scott
contents Natural Language Recommendation (NLRec) generates item suggestions based on the relevance between user-issued NL requests and NL item description passages. Existing NLRec approaches often use Dense Retrieval (DR) to compute item relevance scores from aggregation of inner products between user request embeddings and relevant passage embeddings. However, DR views the request as the sole relevance label, thus leading to a unimodal scoring function centered on the query embedding that is often a weak proxy for query relevance. To better capture the potential multimodal distribution of the relevance scoring function that may arise from complex NLRec data, we propose GPR-LLM that uses Gaussian Process Regression (GPR) with LLM relevance judgments for a subset of candidate passages. Experiments on four NLRec datasets and two LLM backbones demonstrate that GPR-LLM with an RBF kernel, capable of modeling multimodal relevance scoring functions, consistently outperforms simpler unimodal kernels (dot product, cosine similarity), as well as baseline methods including DR, cross-encoder, and pointwise LLM-based relevance scoring by up to 65%. Overall, GPR-LLM provides an efficient and effective approach to NLRec within a minimal LLM labeling budget.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22023
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments
Liu, Yifan
Wen, Qianfeng
Liang, Jiazhou
Zhao, Mark
Cui, Justin
Korikov, Anton
Toroghi, Armin
Kim, Junyoung
Sanner, Scott
Information Retrieval
Natural Language Recommendation (NLRec) generates item suggestions based on the relevance between user-issued NL requests and NL item description passages. Existing NLRec approaches often use Dense Retrieval (DR) to compute item relevance scores from aggregation of inner products between user request embeddings and relevant passage embeddings. However, DR views the request as the sole relevance label, thus leading to a unimodal scoring function centered on the query embedding that is often a weak proxy for query relevance. To better capture the potential multimodal distribution of the relevance scoring function that may arise from complex NLRec data, we propose GPR-LLM that uses Gaussian Process Regression (GPR) with LLM relevance judgments for a subset of candidate passages. Experiments on four NLRec datasets and two LLM backbones demonstrate that GPR-LLM with an RBF kernel, capable of modeling multimodal relevance scoring functions, consistently outperforms simpler unimodal kernels (dot product, cosine similarity), as well as baseline methods including DR, cross-encoder, and pointwise LLM-based relevance scoring by up to 65%. Overall, GPR-LLM provides an efficient and effective approach to NLRec within a minimal LLM labeling budget.
title Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments
topic Information Retrieval
url https://arxiv.org/abs/2510.22023