The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Niousha, Rose, Smith, Samantha Boatright, Akram, Bita, Brusilovsky, Peter, Hellas, Arto, Leinonen, Juho, DeNero, John, Norouzi, Narges
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918487025254400
author Niousha, Rose
Smith, Samantha Boatright
Akram, Bita
Brusilovsky, Peter
Hellas, Arto
Leinonen, Juho
DeNero, John
Norouzi, Narges
author_facet Niousha, Rose
Smith, Samantha Boatright
Akram, Bita
Brusilovsky, Peter
Hellas, Arto
Leinonen, Juho
DeNero, John
Norouzi, Narges
contents Current Artificial Intelligence (AI)-based tutoring systems (AI tutors) are primarily evaluated based on the pedagogical quality of their feedback messages. While important, pedagogy alone is insufficient because it ignores a critical question: what do students actually do with the feedback they receive? We argue that AI tutor evaluation should be extended with a behavioral dimension grounded in student interaction data, which complements pedagogical assessment. We propose an evaluation framework and apply it to 10,235 code submissions with corresponding AI tutor feedback from an introductory undergraduate programming course to measure whether students act on tutor feedback and whether those actions are applied correctly. Using this framework to compare two deployed AI tutors across different semesters in a large-scale introductory computer science course reveals substantial differences in student engagement patterns that are not captured by pedagogy-only evaluation. Moreover, these engagement-based behavioral signals are more strongly associated with student perception of helpful feedback than pedagogical quality alone, providing a more complete and actionable picture of AI tutor performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05648
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
Niousha, Rose
Smith, Samantha Boatright
Akram, Bita
Brusilovsky, Peter
Hellas, Arto
Leinonen, Juho
DeNero, John
Norouzi, Narges
Computers and Society
Artificial Intelligence
Human-Computer Interaction
Current Artificial Intelligence (AI)-based tutoring systems (AI tutors) are primarily evaluated based on the pedagogical quality of their feedback messages. While important, pedagogy alone is insufficient because it ignores a critical question: what do students actually do with the feedback they receive? We argue that AI tutor evaluation should be extended with a behavioral dimension grounded in student interaction data, which complements pedagogical assessment. We propose an evaluation framework and apply it to 10,235 code submissions with corresponding AI tutor feedback from an introductory undergraduate programming course to measure whether students act on tutor feedback and whether those actions are applied correctly. Using this framework to compare two deployed AI tutors across different semesters in a large-scale introductory computer science course reveals substantial differences in student engagement patterns that are not captured by pedagogy-only evaluation. Moreover, these engagement-based behavioral signals are more strongly associated with student perception of helpful feedback than pedagogical quality alone, providing a more complete and actionable picture of AI tutor performance.
title The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
topic Computers and Society
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2605.05648