Optimizing Multi-Stage Language Models for Effective Text Retrieval

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Trung, Quang Hoang, Hoang, Le Trung, Phuc, Nguyen Van Hoang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912170303815680
author Trung, Quang Hoang
Hoang, Le Trung
Phuc, Nguyen Van Hoang
author_facet Trung, Quang Hoang
Hoang, Le Trung
Phuc, Nguyen Van Hoang
contents Efficient text retrieval is critical for applications such as legal document analysis, particularly in specialized contexts like Japanese legal systems. Existing retrieval methods often underperform in such domain-specific scenarios, necessitating tailored approaches. In this paper, we introduce a novel two-phase text retrieval pipeline optimized for Japanese legal datasets. Our method leverages advanced language models to achieve state-of-the-art performance, significantly improving retrieval efficiency and accuracy. To further enhance robustness and adaptability, we incorporate an ensemble model that integrates multiple retrieval strategies, resulting in superior outcomes across diverse tasks. Extensive experiments validate the effectiveness of our approach, demonstrating strong performance on both Japanese legal datasets and widely recognized benchmarks like MS-MARCO. Our work establishes new standards for text retrieval in domain-specific and general contexts, providing a comprehensive solution for addressing complex queries in legal and multilingual environments.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19265
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimizing Multi-Stage Language Models for Effective Text Retrieval
Trung, Quang Hoang
Hoang, Le Trung
Phuc, Nguyen Van Hoang
Information Retrieval
Computation and Language
Machine Learning
Efficient text retrieval is critical for applications such as legal document analysis, particularly in specialized contexts like Japanese legal systems. Existing retrieval methods often underperform in such domain-specific scenarios, necessitating tailored approaches. In this paper, we introduce a novel two-phase text retrieval pipeline optimized for Japanese legal datasets. Our method leverages advanced language models to achieve state-of-the-art performance, significantly improving retrieval efficiency and accuracy. To further enhance robustness and adaptability, we incorporate an ensemble model that integrates multiple retrieval strategies, resulting in superior outcomes across diverse tasks. Extensive experiments validate the effectiveness of our approach, demonstrating strong performance on both Japanese legal datasets and widely recognized benchmarks like MS-MARCO. Our work establishes new standards for text retrieval in domain-specific and general contexts, providing a comprehensive solution for addressing complex queries in legal and multilingual environments.
title Optimizing Multi-Stage Language Models for Effective Text Retrieval
topic Information Retrieval
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.19265