Integrating Vision-Centric Text Understanding for Conversational Recommender Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yuan, Wei, Qiao, Shutong, Chen, Tong, Nguyen, Quoc Viet Hung, Huang, Zi, Yin, Hongzhi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914265547407360
author Yuan, Wei
Qiao, Shutong
Chen, Tong
Nguyen, Quoc Viet Hung
Huang, Zi
Yin, Hongzhi
author_facet Yuan, Wei
Qiao, Shutong
Chen, Tong
Nguyen, Quoc Viet Hung
Huang, Zi
Yin, Hongzhi
contents Conversational Recommender Systems (CRSs) have attracted growing attention for their ability to deliver personalized recommendations through natural language interactions. To more accurately infer user preferences from multi-turn conversations, recent works increasingly expand conversational context (e.g., by incorporating diverse entity information or retrieving related dialogues). While such context enrichment can assist preference modeling, it also introduces longer and more heterogeneous inputs, leading to practical issues such as input length constraints, text style inconsistency, and irrelevant textual noise, thereby raising the demand for stronger language understanding ability. In this paper, we propose STARCRS, a Screen-Text-AwaRe Conversational Recommender System that integrates two complementary text understanding modes: (1) a screen-reading pathway that encodes auxiliary textual information as visual tokens, mimicking skim reading on a screen, and (2) an LLM-based textual pathway that focuses on a limited set of critical content for fine-grained reasoning. We design a knowledge-anchored fusion framework that combines contrastive alignment, cross-attention interaction, and adaptive gating to integrate the two modes for improved preference modeling and response generation. Extensive experiments on two widely used benchmarks demonstrate that STARCRS consistently improves both recommendation accuracy and generated response quality.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13505
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Integrating Vision-Centric Text Understanding for Conversational Recommender Systems
Yuan, Wei
Qiao, Shutong
Chen, Tong
Nguyen, Quoc Viet Hung
Huang, Zi
Yin, Hongzhi
Information Retrieval
Conversational Recommender Systems (CRSs) have attracted growing attention for their ability to deliver personalized recommendations through natural language interactions. To more accurately infer user preferences from multi-turn conversations, recent works increasingly expand conversational context (e.g., by incorporating diverse entity information or retrieving related dialogues). While such context enrichment can assist preference modeling, it also introduces longer and more heterogeneous inputs, leading to practical issues such as input length constraints, text style inconsistency, and irrelevant textual noise, thereby raising the demand for stronger language understanding ability. In this paper, we propose STARCRS, a Screen-Text-AwaRe Conversational Recommender System that integrates two complementary text understanding modes: (1) a screen-reading pathway that encodes auxiliary textual information as visual tokens, mimicking skim reading on a screen, and (2) an LLM-based textual pathway that focuses on a limited set of critical content for fine-grained reasoning. We design a knowledge-anchored fusion framework that combines contrastive alignment, cross-attention interaction, and adaptive gating to integrate the two modes for improved preference modeling and response generation. Extensive experiments on two widely used benchmarks demonstrate that STARCRS consistently improves both recommendation accuracy and generated response quality.
title Integrating Vision-Centric Text Understanding for Conversational Recommender Systems
topic Information Retrieval
url https://arxiv.org/abs/2601.13505