Position: Thematic Analysis of Unstructured Clinical Transcripts with Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908562632998912 |
|---|---|
| author | Yi, Seungjun Nguyen, Joakim Lim, Terence Well, Andrew Skrovan, Joseph Beri, Mehak Lee, YongGeon Radhakrishnan, Kavita Leqi, Liu Markey, Mia Ding, Ying |
| author_facet | Yi, Seungjun Nguyen, Joakim Lim, Terence Well, Andrew Skrovan, Joseph Beri, Mehak Lee, YongGeon Radhakrishnan, Kavita Leqi, Liu Markey, Mia Ding, Ying |
| contents | This position paper examines how large language models (LLMs) can support thematic analysis of unstructured clinical transcripts, a widely used but resource-intensive method for uncovering patterns in patient and provider narratives. We conducted a systematic review of recent studies applying LLMs to thematic analysis, complemented by an interview with a practicing clinician. Our findings reveal that current approaches remain fragmented across multiple dimensions including types of thematic analysis, datasets, prompting strategies and models used, most notably in evaluation. Existing evaluation methods vary widely (from qualitative expert review to automatic similarity metrics), hindering progress and preventing meaningful benchmarking across studies. We argue that establishing standardized evaluation practices is critical for advancing the field. To this end, we propose an evaluation framework centered on three dimensions: validity, reliability, and interpretability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_14597 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Position: Thematic Analysis of Unstructured Clinical Transcripts with Large Language Models Yi, Seungjun Nguyen, Joakim Lim, Terence Well, Andrew Skrovan, Joseph Beri, Mehak Lee, YongGeon Radhakrishnan, Kavita Leqi, Liu Markey, Mia Ding, Ying Computation and Language This position paper examines how large language models (LLMs) can support thematic analysis of unstructured clinical transcripts, a widely used but resource-intensive method for uncovering patterns in patient and provider narratives. We conducted a systematic review of recent studies applying LLMs to thematic analysis, complemented by an interview with a practicing clinician. Our findings reveal that current approaches remain fragmented across multiple dimensions including types of thematic analysis, datasets, prompting strategies and models used, most notably in evaluation. Existing evaluation methods vary widely (from qualitative expert review to automatic similarity metrics), hindering progress and preventing meaningful benchmarking across studies. We argue that establishing standardized evaluation practices is critical for advancing the field. To this end, we propose an evaluation framework centered on three dimensions: validity, reliability, and interpretability. |
| title | Position: Thematic Analysis of Unstructured Clinical Transcripts with Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2509.14597 |