From Logs to Language: Learning Optimal Verbalization for LLM-Based Recommendation at Industry Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Yucheng, Li, Ying, Wang, Yu, Feng, Yesu, Rao, Arjun, Houthooft, Rein, Sehgal, Shradha, Wang, Jin, Zhen, Hao, Liu, Ninghao, Baltrunas, Linas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911527682965504
author Shi, Yucheng
Li, Ying
Wang, Yu
Feng, Yesu
Rao, Arjun
Houthooft, Rein
Sehgal, Shradha
Wang, Jin
Zhen, Hao
Liu, Ninghao
Baltrunas, Linas
author_facet Shi, Yucheng
Li, Ying
Wang, Yu
Feng, Yesu
Rao, Arjun
Houthooft, Rein
Sehgal, Shradha
Wang, Jin
Zhen, Hao
Liu, Ninghao
Baltrunas, Linas
contents Large language models (LLMs) are promising backbones for generative recommender systems, yet a key challenge remains underexplored: verbalization, i.e., converting structured user interaction logs into effective natural language inputs. Existing methods rely on rigid templates that simply concatenate fields, yielding suboptimal representations for recommendation. We propose a data-centric framework that learns verbalization for LLM-based recommendation. Using reinforcement learning, a verbalization agent transforms raw interaction histories into optimized textual contexts, with recommendation accuracy as the training signal. This agent learns to filter noise, incorporate relevant metadata, and reorganize information to improve downstream predictions. Experiments on a large-scale industrial streaming dataset from Netflix show that learned verbalization delivers up to 93% relative improvement in discovery item recommendation accuracy over template-based baselines. Further analysis reveals emergent strategies such as user interest summarization, noise removal, and syntax normalization, offering insights into effective context construction for LLM-based recommender systems.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20558
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Logs to Language: Learning Optimal Verbalization for LLM-Based Recommendation at Industry Scale
Shi, Yucheng
Li, Ying
Wang, Yu
Feng, Yesu
Rao, Arjun
Houthooft, Rein
Sehgal, Shradha
Wang, Jin
Zhen, Hao
Liu, Ninghao
Baltrunas, Linas
Artificial Intelligence
Information Retrieval
Large language models (LLMs) are promising backbones for generative recommender systems, yet a key challenge remains underexplored: verbalization, i.e., converting structured user interaction logs into effective natural language inputs. Existing methods rely on rigid templates that simply concatenate fields, yielding suboptimal representations for recommendation. We propose a data-centric framework that learns verbalization for LLM-based recommendation. Using reinforcement learning, a verbalization agent transforms raw interaction histories into optimized textual contexts, with recommendation accuracy as the training signal. This agent learns to filter noise, incorporate relevant metadata, and reorganize information to improve downstream predictions. Experiments on a large-scale industrial streaming dataset from Netflix show that learned verbalization delivers up to 93% relative improvement in discovery item recommendation accuracy over template-based baselines. Further analysis reveals emergent strategies such as user interest summarization, noise removal, and syntax normalization, offering insights into effective context construction for LLM-based recommender systems.
title From Logs to Language: Learning Optimal Verbalization for LLM-Based Recommendation at Industry Scale
topic Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2602.20558