ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Singh, Ishneet Sukhvinder, Aggarwal, Ritvik, Allahverdiyev, Ibrahim, Taha, Muhammad, Akalin, Aslihan, Zhu, Kevin, O'Brien, Sean
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910917088772096
author Singh, Ishneet Sukhvinder
Aggarwal, Ritvik
Allahverdiyev, Ibrahim
Taha, Muhammad
Akalin, Aslihan
Zhu, Kevin
O'Brien, Sean
author_facet Singh, Ishneet Sukhvinder
Aggarwal, Ritvik
Allahverdiyev, Ibrahim
Taha, Muhammad
Akalin, Aslihan
Zhu, Kevin
O'Brien, Sean
contents Retrieval-Augmented Generation (RAG) systems using large language models (LLMs) often generate inaccurate responses due to the retrieval of irrelevant or loosely related information. Existing methods, which operate at the document level, fail to effectively filter out such content. We propose LLM-driven chunk filtering, ChunkRAG, a framework that enhances RAG systems by evaluating and filtering retrieved information at the chunk level. Our approach employs semantic chunking to divide documents into coherent sections and utilizes LLM-based relevance scoring to assess each chunk's alignment with the user's query. By filtering out less pertinent chunks before the generation phase, we significantly reduce hallucinations and improve factual accuracy. Experiments show that our method outperforms existing RAG models, achieving higher accuracy on tasks requiring precise information retrieval. This advancement enhances the reliability of RAG systems, making them particularly beneficial for applications like fact-checking and multi-hop reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19572
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
Singh, Ishneet Sukhvinder
Aggarwal, Ritvik
Allahverdiyev, Ibrahim
Taha, Muhammad
Akalin, Aslihan
Zhu, Kevin
O'Brien, Sean
Computation and Language
Retrieval-Augmented Generation (RAG) systems using large language models (LLMs) often generate inaccurate responses due to the retrieval of irrelevant or loosely related information. Existing methods, which operate at the document level, fail to effectively filter out such content. We propose LLM-driven chunk filtering, ChunkRAG, a framework that enhances RAG systems by evaluating and filtering retrieved information at the chunk level. Our approach employs semantic chunking to divide documents into coherent sections and utilizes LLM-based relevance scoring to assess each chunk's alignment with the user's query. By filtering out less pertinent chunks before the generation phase, we significantly reduce hallucinations and improve factual accuracy. Experiments show that our method outperforms existing RAG models, achieving higher accuracy on tasks requiring precise information retrieval. This advancement enhances the reliability of RAG systems, making them particularly beneficial for applications like fact-checking and multi-hop reasoning.
title ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
topic Computation and Language
url https://arxiv.org/abs/2410.19572