Toward LLM-Supported Automated Assessment of Critical Thinking Subskills

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Peczuh, Marisa C., Kumar, Nischal Ashok, Baker, Ryan, Lehman, Blair, Eisenberg, Danielle, Mills, Caitlin, Wittawatolarn, Payu, Naskar, Kushaan, Chebrolu, Keerthi, Nashi, Sudhip, Young, Cadence, Liu, Brayden, Lachman, Sherry, Lan, Andrew
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914337975697408
author Peczuh, Marisa C.
Kumar, Nischal Ashok
Baker, Ryan
Lehman, Blair
Eisenberg, Danielle
Mills, Caitlin
Wittawatolarn, Payu
Naskar, Kushaan
Chebrolu, Keerthi
Nashi, Sudhip
Young, Cadence
Liu, Brayden
Lachman, Sherry
Lan, Andrew
author_facet Peczuh, Marisa C.
Kumar, Nischal Ashok
Baker, Ryan
Lehman, Blair
Eisenberg, Danielle
Mills, Caitlin
Wittawatolarn, Payu
Naskar, Kushaan
Chebrolu, Keerthi
Nashi, Sudhip
Young, Cadence
Liu, Brayden
Lachman, Sherry
Lan, Andrew
contents As the world becomes increasingly saturated with AI-generated content, disinformation, and algorithmic persuasion, critical thinking - the capacity to evaluate evidence, detect unreliable claims, and exercise independent judgment - is becoming a defining human skill. Developing critical thinking skills through timely assessment and feedback is crucial; however, there has not been extensive work in educational data mining on defining, measuring, and supporting critical thinking. In this paper, we investigate the feasibility of measuring "subskills" that underlie critical thinking. We ground our work in an authentic task where students operationalize critical thinking by writing argumentative essays. We developed a coding rubric based on an established skills progression and completed human coding for a corpus of student essays. We then evaluated three distinct approaches to automated scoring: zero-shot prompting, few-shot prompting, and supervised fine-tuning, implemented across three large language models (GPT-5, Llama 3.1 8B, and ModernBERT). Fine-tuning Llama 3.1 8B achieved the best results and demonstrated particular strength on subskills with highly separable proficiency levels with balanced labels across levels, while lower performance was observed for subskills that required detection of subtle distinctions between proficiency levels or imbalanced labels. Our exploratory work represents an initial step toward scalable assessment of critical thinking skills across authentic educational contexts. Future research should continue to combine automated critical thinking assessment with human validation to more accurately detect and measure dynamic, higher-order thinking skills.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
Peczuh, Marisa C.
Kumar, Nischal Ashok
Baker, Ryan
Lehman, Blair
Eisenberg, Danielle
Mills, Caitlin
Wittawatolarn, Payu
Naskar, Kushaan
Chebrolu, Keerthi
Nashi, Sudhip
Young, Cadence
Liu, Brayden
Lachman, Sherry
Lan, Andrew
Computers and Society
Computation and Language
Machine Learning
As the world becomes increasingly saturated with AI-generated content, disinformation, and algorithmic persuasion, critical thinking - the capacity to evaluate evidence, detect unreliable claims, and exercise independent judgment - is becoming a defining human skill. Developing critical thinking skills through timely assessment and feedback is crucial; however, there has not been extensive work in educational data mining on defining, measuring, and supporting critical thinking. In this paper, we investigate the feasibility of measuring "subskills" that underlie critical thinking. We ground our work in an authentic task where students operationalize critical thinking by writing argumentative essays. We developed a coding rubric based on an established skills progression and completed human coding for a corpus of student essays. We then evaluated three distinct approaches to automated scoring: zero-shot prompting, few-shot prompting, and supervised fine-tuning, implemented across three large language models (GPT-5, Llama 3.1 8B, and ModernBERT). Fine-tuning Llama 3.1 8B achieved the best results and demonstrated particular strength on subskills with highly separable proficiency levels with balanced labels across levels, while lower performance was observed for subskills that required detection of subtle distinctions between proficiency levels or imbalanced labels. Our exploratory work represents an initial step toward scalable assessment of critical thinking skills across authentic educational contexts. Future research should continue to combine automated critical thinking assessment with human validation to more accurately detect and measure dynamic, higher-order thinking skills.
title Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
topic Computers and Society
Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.12915