Guardado en:
Detalles Bibliográficos
Autores principales: Hao, Jiangang, Cui, Wenju, Kyllonen, Patrick, Kerzabi, Emily
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2510.20584
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909053286875136
author Hao, Jiangang
Cui, Wenju
Kyllonen, Patrick
Kerzabi, Emily
author_facet Hao, Jiangang
Cui, Wenju
Kyllonen, Patrick
Kerzabi, Emily
contents Assessing communication and collaboration at scale depends on a labor-intensive task of coding communication data into categories according to different frameworks. Prior research has established that ChatGPT can be directly instructed with coding rubrics to code the communication data and achieves accuracy comparable to human raters. However, whether the coding from ChatGPT or similar AI technology perform consistently across different demographic groups, such as gender and race, remains unclear. To address this gap, we introduce three checks for evaluating subgroup consistency in LLM-based coding by adapting an existing framework from the automated scoring literature. Using a typical collaborative problem-solving coding framework and data from three types of collaborative tasks, we examine ChatGPT-based coding performance across gender and racial/ethnic groups. Our results show that ChatGPT-based coding perform consistently in the same way as human raters across gender or racial/ethnic groups, demonstrating the possibility of its use in large-scale assessments of collaboration and communication.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20584
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automated Coding of Communication Data Using ChatGPT: Consistency Across Subgroups
Hao, Jiangang
Cui, Wenju
Kyllonen, Patrick
Kerzabi, Emily
Computation and Language
Artificial Intelligence
Assessing communication and collaboration at scale depends on a labor-intensive task of coding communication data into categories according to different frameworks. Prior research has established that ChatGPT can be directly instructed with coding rubrics to code the communication data and achieves accuracy comparable to human raters. However, whether the coding from ChatGPT or similar AI technology perform consistently across different demographic groups, such as gender and race, remains unclear. To address this gap, we introduce three checks for evaluating subgroup consistency in LLM-based coding by adapting an existing framework from the automated scoring literature. Using a typical collaborative problem-solving coding framework and data from three types of collaborative tasks, we examine ChatGPT-based coding performance across gender and racial/ethnic groups. Our results show that ChatGPT-based coding perform consistently in the same way as human raters across gender or racial/ethnic groups, demonstrating the possibility of its use in large-scale assessments of collaboration and communication.
title Automated Coding of Communication Data Using ChatGPT: Consistency Across Subgroups
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.20584