Talking to Data: Designing Smart Assistants for Humanities Databases

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sergeev, Alexander, Goloviznina, Valeriya, Melnichenko, Mikhail, Kotelnikov, Evgeny
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915317537570816
author Sergeev, Alexander
Goloviznina, Valeriya
Melnichenko, Mikhail
Kotelnikov, Evgeny
author_facet Sergeev, Alexander
Goloviznina, Valeriya
Melnichenko, Mikhail
Kotelnikov, Evgeny
contents Access to humanities research databases is often hindered by the limitations of traditional interaction formats, particularly in the methods of searching and response generation. This study introduces an LLM-based smart assistant designed to facilitate natural language communication with digital humanities data. The assistant, developed in a chatbot format, leverages the RAG approach and integrates state-of-the-art technologies such as hybrid search, automatic query generation, text-to-SQL filtering, semantic database search, and hyperlink insertion. To evaluate the effectiveness of the system, experiments were conducted to assess the response quality of various language models. The testing was based on the Prozhito digital archive, which contains diary entries from predominantly Russian-speaking individuals who lived in the 20th century. The chatbot is tailored to support anthropology and history researchers, as well as non-specialist users with an interest in the field, without requiring prior technical training. By enabling researchers to query complex databases with natural language, this tool aims to enhance accessibility and efficiency in humanities research. The study highlights the potential of Large Language Models to transform the way researchers and the public interact with digital archives, making them more intuitive and inclusive. Additional materials are presented in GitHub repository: https://github.com/alekosus/talking-to-data-intersys2025.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00986
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Talking to Data: Designing Smart Assistants for Humanities Databases
Sergeev, Alexander
Goloviznina, Valeriya
Melnichenko, Mikhail
Kotelnikov, Evgeny
Computation and Language
Access to humanities research databases is often hindered by the limitations of traditional interaction formats, particularly in the methods of searching and response generation. This study introduces an LLM-based smart assistant designed to facilitate natural language communication with digital humanities data. The assistant, developed in a chatbot format, leverages the RAG approach and integrates state-of-the-art technologies such as hybrid search, automatic query generation, text-to-SQL filtering, semantic database search, and hyperlink insertion. To evaluate the effectiveness of the system, experiments were conducted to assess the response quality of various language models. The testing was based on the Prozhito digital archive, which contains diary entries from predominantly Russian-speaking individuals who lived in the 20th century. The chatbot is tailored to support anthropology and history researchers, as well as non-specialist users with an interest in the field, without requiring prior technical training. By enabling researchers to query complex databases with natural language, this tool aims to enhance accessibility and efficiency in humanities research. The study highlights the potential of Large Language Models to transform the way researchers and the public interact with digital archives, making them more intuitive and inclusive. Additional materials are presented in GitHub repository: https://github.com/alekosus/talking-to-data-intersys2025.
title Talking to Data: Designing Smart Assistants for Humanities Databases
topic Computation and Language
url https://arxiv.org/abs/2506.00986