Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mao, Qiang, Qin, Han, Neary, Robert, Wang, Charles, Wei, Fusheng, Zhang, Jianping, Huber-Fliflet, Nathaniel
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2512.08078
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912754983501824
author Mao, Qiang
Qin, Han
Neary, Robert
Wang, Charles
Wei, Fusheng
Zhang, Jianping
Huber-Fliflet, Nathaniel
author_facet Mao, Qiang
Qin, Han
Neary, Robert
Wang, Charles
Wei, Fusheng
Zhang, Jianping
Huber-Fliflet, Nathaniel
contents Increasingly, attorneys are interested in moving beyond keyword and semantic search to improve the efficiency of how they find key information during a document review task. Large language models (LLMs) are now seen as tools that attorneys can use to ask natural language questions of their data during document review to receive accurate and concise answers. This study evaluates retrieval strategies within Microsoft Azure's Retrieval-Augmented Generation (RAG) framework to identify effective approaches for Early Case Assessment (ECA) in eDiscovery. During ECA, legal teams analyze data at the outset of a matter to gain a general understanding of the data and attempt to determine key facts and risks before beginning full-scale review. In this paper, we compare the performance of Azure AI Search's keyword, semantic, vector, hybrid, and hybrid-semantic retrieval methods. We then present the accuracy, relevance, and consistency of each method's AI-generated responses. Legal practitioners can use the results of this study to enhance how they select RAG configurations in the future.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08078
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Comparative Study of Retrieval Methods in Azure AI Search
Mao, Qiang
Qin, Han
Neary, Robert
Wang, Charles
Wei, Fusheng
Zhang, Jianping
Huber-Fliflet, Nathaniel
Information Retrieval
Increasingly, attorneys are interested in moving beyond keyword and semantic search to improve the efficiency of how they find key information during a document review task. Large language models (LLMs) are now seen as tools that attorneys can use to ask natural language questions of their data during document review to receive accurate and concise answers. This study evaluates retrieval strategies within Microsoft Azure's Retrieval-Augmented Generation (RAG) framework to identify effective approaches for Early Case Assessment (ECA) in eDiscovery. During ECA, legal teams analyze data at the outset of a matter to gain a general understanding of the data and attempt to determine key facts and risks before beginning full-scale review. In this paper, we compare the performance of Azure AI Search's keyword, semantic, vector, hybrid, and hybrid-semantic retrieval methods. We then present the accuracy, relevance, and consistency of each method's AI-generated responses. Legal practitioners can use the results of this study to enhance how they select RAG configurations in the future.
title A Comparative Study of Retrieval Methods in Azure AI Search
topic Information Retrieval
url https://arxiv.org/abs/2512.08078