From keywords to semantics: Perceptions of large language models in data discovery
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915530534813696 |
|---|---|
| author | Halstead, Maura E Green, Mark A. Jay, Caroline Kingston, Richard Topping, David Singleton, Alexander |
| author_facet | Halstead, Maura E Green, Mark A. Jay, Caroline Kingston, Richard Topping, David Singleton, Alexander |
| contents | Current approaches to data discovery match keywords between metadata and queries. This matching requires researchers to know the exact wording that other researchers previously used, creating a challenging process that could lead to missing relevant data. Large Language Models (LLMs) could enhance data discovery by removing this requirement and allowing researchers to ask questions with natural language. However, we do not currently know if researchers would accept LLMs for data discovery. Using a human-centered artificial intelligence (HCAI) focus, we ran focus groups (N = 27) to understand researchers' perspectives towards LLMs for data discovery. Our conceptual model shows that the potential benefits are not enough for researchers to use LLMs instead of current technology. Barriers prevent researchers from fully accepting LLMs, but features around transparency could overcome them. Using our model will allow developers to incorporate features that result in an increased acceptance of LLMs for data discovery. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_01473 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | From keywords to semantics: Perceptions of large language models in data discovery Halstead, Maura E Green, Mark A. Jay, Caroline Kingston, Richard Topping, David Singleton, Alexander Human-Computer Interaction Artificial Intelligence Current approaches to data discovery match keywords between metadata and queries. This matching requires researchers to know the exact wording that other researchers previously used, creating a challenging process that could lead to missing relevant data. Large Language Models (LLMs) could enhance data discovery by removing this requirement and allowing researchers to ask questions with natural language. However, we do not currently know if researchers would accept LLMs for data discovery. Using a human-centered artificial intelligence (HCAI) focus, we ran focus groups (N = 27) to understand researchers' perspectives towards LLMs for data discovery. Our conceptual model shows that the potential benefits are not enough for researchers to use LLMs instead of current technology. Barriers prevent researchers from fully accepting LLMs, but features around transparency could overcome them. Using our model will allow developers to incorporate features that result in an increased acceptance of LLMs for data discovery. |
| title | From keywords to semantics: Perceptions of large language models in data discovery |
| topic | Human-Computer Interaction Artificial Intelligence |
| url | https://arxiv.org/abs/2510.01473 |