Enhancing Question Answering Precision with Optimized Vector Retrieval and Instructions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Lixiao, Xu, Mengyang, Ke, Weimao
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929575028588544
author Yang, Lixiao
Xu, Mengyang
Ke, Weimao
author_facet Yang, Lixiao
Xu, Mengyang
Ke, Weimao
contents Question-answering (QA) is an important application of Information Retrieval (IR) and language models, and the latest trend is toward pre-trained large neural networks with embedding parameters. Augmenting QA performances with these LLMs requires intensive computational resources for fine-tuning. We propose an innovative approach to improve QA task performances by integrating optimized vector retrievals and instruction methodologies. Based on retrieval augmentation, the process involves document embedding, vector retrieval, and context construction for optimal QA results. We experiment with different combinations of text segmentation techniques and similarity functions, and analyze their impacts on QA performances. Results show that the model with a small chunk size of 100 without any overlap of the chunks achieves the best result and outperforms the models based on semantic segmentation using sentences. We discuss related QA examples and offer insight into how model performances are improved within the two-stage framework.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Question Answering Precision with Optimized Vector Retrieval and Instructions
Yang, Lixiao
Xu, Mengyang
Ke, Weimao
Information Retrieval
Computation and Language
Machine Learning
Question-answering (QA) is an important application of Information Retrieval (IR) and language models, and the latest trend is toward pre-trained large neural networks with embedding parameters. Augmenting QA performances with these LLMs requires intensive computational resources for fine-tuning. We propose an innovative approach to improve QA task performances by integrating optimized vector retrievals and instruction methodologies. Based on retrieval augmentation, the process involves document embedding, vector retrieval, and context construction for optimal QA results. We experiment with different combinations of text segmentation techniques and similarity functions, and analyze their impacts on QA performances. Results show that the model with a small chunk size of 100 without any overlap of the chunks achieves the best result and outperforms the models based on semantic segmentation using sentences. We discuss related QA examples and offer insight into how model performances are improved within the two-stage framework.
title Enhancing Question Answering Precision with Optimized Vector Retrieval and Instructions
topic Information Retrieval
Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.01039