A Comparative Study of Retrieval Methods in Azure AI Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Qiang, Qin, Han, Neary, Robert, Wang, Charles, Wei, Fusheng, Zhang, Jianping, Huber-Fliflet, Nathaniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912754983501824
author Mao, Qiang
Qin, Han
Neary, Robert
Wang, Charles
Wei, Fusheng
Zhang, Jianping
Huber-Fliflet, Nathaniel
author_facet Mao, Qiang
Qin, Han
Neary, Robert
Wang, Charles
Wei, Fusheng
Zhang, Jianping
Huber-Fliflet, Nathaniel
contents Increasingly, attorneys are interested in moving beyond keyword and semantic search to improve the efficiency of how they find key information during a document review task. Large language models (LLMs) are now seen as tools that attorneys can use to ask natural language questions of their data during document review to receive accurate and concise answers. This study evaluates retrieval strategies within Microsoft Azure's Retrieval-Augmented Generation (RAG) framework to identify effective approaches for Early Case Assessment (ECA) in eDiscovery. During ECA, legal teams analyze data at the outset of a matter to gain a general understanding of the data and attempt to determine key facts and risks before beginning full-scale review. In this paper, we compare the performance of Azure AI Search's keyword, semantic, vector, hybrid, and hybrid-semantic retrieval methods. We then present the accuracy, relevance, and consistency of each method's AI-generated responses. Legal practitioners can use the results of this study to enhance how they select RAG configurations in the future.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08078
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Comparative Study of Retrieval Methods in Azure AI Search
Mao, Qiang
Qin, Han
Neary, Robert
Wang, Charles
Wei, Fusheng
Zhang, Jianping
Huber-Fliflet, Nathaniel
Information Retrieval
Increasingly, attorneys are interested in moving beyond keyword and semantic search to improve the efficiency of how they find key information during a document review task. Large language models (LLMs) are now seen as tools that attorneys can use to ask natural language questions of their data during document review to receive accurate and concise answers. This study evaluates retrieval strategies within Microsoft Azure's Retrieval-Augmented Generation (RAG) framework to identify effective approaches for Early Case Assessment (ECA) in eDiscovery. During ECA, legal teams analyze data at the outset of a matter to gain a general understanding of the data and attempt to determine key facts and risks before beginning full-scale review. In this paper, we compare the performance of Azure AI Search's keyword, semantic, vector, hybrid, and hybrid-semantic retrieval methods. We then present the accuracy, relevance, and consistency of each method's AI-generated responses. Legal practitioners can use the results of this study to enhance how they select RAG configurations in the future.
title A Comparative Study of Retrieval Methods in Azure AI Search
topic Information Retrieval
url https://arxiv.org/abs/2512.08078