Leveraging Translation For Optimal Recall: Tailoring LLM Personalization With User Profiles

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ravichandran, Karthik, Gomasta, Sarmistha Sarna
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914687704104960
author Ravichandran, Karthik
Gomasta, Sarmistha Sarna
author_facet Ravichandran, Karthik
Gomasta, Sarmistha Sarna
contents This paper explores a novel technique for improving recall in cross-language information retrieval (CLIR) systems using iterative query refinement grounded in the user's lexical-semantic space. The proposed methodology combines multi-level translation, semantic embedding-based expansion, and user profile-centered augmentation to address the challenge of matching variance between user queries and relevant documents. Through an initial BM25 retrieval, translation into intermediate languages, embedding lookup of similar terms, and iterative re-ranking, the technique aims to expand the scope of potentially relevant results personalized to the individual user. Comparative experiments on news and Twitter datasets demonstrate superior performance over baseline BM25 ranking for the proposed approach across ROUGE metrics. The translation methodology also showed maintained semantic accuracy through the multi-step process. This personalized CLIR framework paves the path for improved context-aware retrieval attentive to the nuances of user language.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13500
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Translation For Optimal Recall: Tailoring LLM Personalization With User Profiles
Ravichandran, Karthik
Gomasta, Sarmistha Sarna
Information Retrieval
Computation and Language
F.2.2; I.2.7
This paper explores a novel technique for improving recall in cross-language information retrieval (CLIR) systems using iterative query refinement grounded in the user's lexical-semantic space. The proposed methodology combines multi-level translation, semantic embedding-based expansion, and user profile-centered augmentation to address the challenge of matching variance between user queries and relevant documents. Through an initial BM25 retrieval, translation into intermediate languages, embedding lookup of similar terms, and iterative re-ranking, the technique aims to expand the scope of potentially relevant results personalized to the individual user. Comparative experiments on news and Twitter datasets demonstrate superior performance over baseline BM25 ranking for the proposed approach across ROUGE metrics. The translation methodology also showed maintained semantic accuracy through the multi-step process. This personalized CLIR framework paves the path for improved context-aware retrieval attentive to the nuances of user language.
title Leveraging Translation For Optimal Recall: Tailoring LLM Personalization With User Profiles
topic Information Retrieval
Computation and Language
F.2.2; I.2.7
url https://arxiv.org/abs/2402.13500