From keywords to semantics: Perceptions of large language models in data discovery

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Halstead, Maura E, Green, Mark A., Jay, Caroline, Kingston, Richard, Topping, David, Singleton, Alexander
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915530534813696
author Halstead, Maura E
Green, Mark A.
Jay, Caroline
Kingston, Richard
Topping, David
Singleton, Alexander
author_facet Halstead, Maura E
Green, Mark A.
Jay, Caroline
Kingston, Richard
Topping, David
Singleton, Alexander
contents Current approaches to data discovery match keywords between metadata and queries. This matching requires researchers to know the exact wording that other researchers previously used, creating a challenging process that could lead to missing relevant data. Large Language Models (LLMs) could enhance data discovery by removing this requirement and allowing researchers to ask questions with natural language. However, we do not currently know if researchers would accept LLMs for data discovery. Using a human-centered artificial intelligence (HCAI) focus, we ran focus groups (N = 27) to understand researchers' perspectives towards LLMs for data discovery. Our conceptual model shows that the potential benefits are not enough for researchers to use LLMs instead of current technology. Barriers prevent researchers from fully accepting LLMs, but features around transparency could overcome them. Using our model will allow developers to incorporate features that result in an increased acceptance of LLMs for data discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From keywords to semantics: Perceptions of large language models in data discovery
Halstead, Maura E
Green, Mark A.
Jay, Caroline
Kingston, Richard
Topping, David
Singleton, Alexander
Human-Computer Interaction
Artificial Intelligence
Current approaches to data discovery match keywords between metadata and queries. This matching requires researchers to know the exact wording that other researchers previously used, creating a challenging process that could lead to missing relevant data. Large Language Models (LLMs) could enhance data discovery by removing this requirement and allowing researchers to ask questions with natural language. However, we do not currently know if researchers would accept LLMs for data discovery. Using a human-centered artificial intelligence (HCAI) focus, we ran focus groups (N = 27) to understand researchers' perspectives towards LLMs for data discovery. Our conceptual model shows that the potential benefits are not enough for researchers to use LLMs instead of current technology. Barriers prevent researchers from fully accepting LLMs, but features around transparency could overcome them. Using our model will allow developers to incorporate features that result in an increased acceptance of LLMs for data discovery.
title From keywords to semantics: Perceptions of large language models in data discovery
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2510.01473