Data Therapist: Eliciting Domain Knowledge from Subject Matter Experts Using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shin, Sungbok, Jeon, Hyeon, Hong, Sanghyun, Elmqvist, Niklas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908621748568064
author Shin, Sungbok
Jeon, Hyeon
Hong, Sanghyun
Elmqvist, Niklas
author_facet Shin, Sungbok
Jeon, Hyeon
Hong, Sanghyun
Elmqvist, Niklas
contents Effective data visualization requires not only technical proficiency but also a deep understanding of the domain-specific context in which data exists. This context often includes tacit knowledge about data provenance, quality, and intended use, which is rarely explicit in the dataset itself. Motivated by growing demands to surface tacit knowledge, we present the Data Therapist, a web-based system that helps domain experts externalize such implicit knowledge through a mixed-initiative process combining iterative Q&A with interactive annotation. Powered by a large language model, the system automatically analyzes user-supplied datasets, prompts users with targeted questions, and supports annotation at varying levels of granularity. The resulting structured knowledge base can inform both human and automated visualization design. A qualitative study with expert pairs from Accounting, Political Science, and Computer Security revealed recurring patterns in how expert reason about their data and highlighted opportunities for AI support to enhance visualization design.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00455
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Therapist: Eliciting Domain Knowledge from Subject Matter Experts Using Large Language Models
Shin, Sungbok
Jeon, Hyeon
Hong, Sanghyun
Elmqvist, Niklas
Human-Computer Interaction
Artificial Intelligence
Effective data visualization requires not only technical proficiency but also a deep understanding of the domain-specific context in which data exists. This context often includes tacit knowledge about data provenance, quality, and intended use, which is rarely explicit in the dataset itself. Motivated by growing demands to surface tacit knowledge, we present the Data Therapist, a web-based system that helps domain experts externalize such implicit knowledge through a mixed-initiative process combining iterative Q&A with interactive annotation. Powered by a large language model, the system automatically analyzes user-supplied datasets, prompts users with targeted questions, and supports annotation at varying levels of granularity. The resulting structured knowledge base can inform both human and automated visualization design. A qualitative study with expert pairs from Accounting, Political Science, and Computer Security revealed recurring patterns in how expert reason about their data and highlighted opportunities for AI support to enhance visualization design.
title Data Therapist: Eliciting Domain Knowledge from Subject Matter Experts Using Large Language Models
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2505.00455