Saved in:
Bibliographic Details
Main Authors: Yang, Hang, Guo, Jing, Qi, Jianchuan, Xie, Jinliang, Zhang, Si, Yang, Siqi, Li, Nan, Xu, Ming
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2405.03989
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913344619806720
author Yang, Hang
Guo, Jing
Qi, Jianchuan
Xie, Jinliang
Zhang, Si
Yang, Siqi
Li, Nan
Xu, Ming
author_facet Yang, Hang
Guo, Jing
Qi, Jianchuan
Xie, Jinliang
Zhang, Si
Yang, Siqi
Li, Nan
Xu, Ming
contents This paper presents a novel method for parsing and vectorizing semi-structured data to enhance the functionality of Retrieval-Augmented Generation (RAG) within Large Language Models (LLMs). We developed a comprehensive pipeline for converting various data formats into .docx, enabling efficient parsing and structured data extraction. The core of our methodology involves the construction of a vector database using Pinecone, which integrates seamlessly with LLMs to provide accurate, context-specific responses, particularly in environmental management and wastewater treatment operations. Through rigorous testing with both English and Chinese texts in diverse document formats, our results demonstrate a marked improvement in the precision and reliability of LLMs outputs. The RAG-enhanced models displayed enhanced ability to generate contextually rich and technically accurate responses, underscoring the potential of vector knowledge bases in significantly boosting the performance of LLMs in specialized domains. This research not only illustrates the effectiveness of our method but also highlights its potential to revolutionize data processing and analysis in environmental sciences, setting a precedent for future advancements in AI-driven applications. Our code is available at https://github.com/linancn/TianGong-AI-Unstructure.git.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03989
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Method for Parsing and Vectorization of Semi-structured Data used in Retrieval Augmented Generation
Yang, Hang
Guo, Jing
Qi, Jianchuan
Xie, Jinliang
Zhang, Si
Yang, Siqi
Li, Nan
Xu, Ming
Databases
This paper presents a novel method for parsing and vectorizing semi-structured data to enhance the functionality of Retrieval-Augmented Generation (RAG) within Large Language Models (LLMs). We developed a comprehensive pipeline for converting various data formats into .docx, enabling efficient parsing and structured data extraction. The core of our methodology involves the construction of a vector database using Pinecone, which integrates seamlessly with LLMs to provide accurate, context-specific responses, particularly in environmental management and wastewater treatment operations. Through rigorous testing with both English and Chinese texts in diverse document formats, our results demonstrate a marked improvement in the precision and reliability of LLMs outputs. The RAG-enhanced models displayed enhanced ability to generate contextually rich and technically accurate responses, underscoring the potential of vector knowledge bases in significantly boosting the performance of LLMs in specialized domains. This research not only illustrates the effectiveness of our method but also highlights its potential to revolutionize data processing and analysis in environmental sciences, setting a precedent for future advancements in AI-driven applications. Our code is available at https://github.com/linancn/TianGong-AI-Unstructure.git.
title A Method for Parsing and Vectorization of Semi-structured Data used in Retrieval Augmented Generation
topic Databases
url https://arxiv.org/abs/2405.03989