The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
Fuente:
arXiv
Saved in:
| Main Authors: | Niousha, Rose, Smith, Samantha Boatright, Akram, Bita, Brusilovsky, Peter, Hellas, Arto, Leinonen, Juho, DeNero, John, Norouzi, Narges |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Personalized Worked Example Generation from Student Code Submissions Using Pattern-based Knowledge Components
by: Pitts, Griffin, et al.
Published: (2026)
by: Pitts, Griffin, et al.
Published: (2026)
Retrieval-Augmented Tutoring for Algorithm Tracing and Problem-Solving in AI Education
by: Jain, Mragisha, et al.
Published: (2026)
by: Jain, Mragisha, et al.
Published: (2026)
The Effects of Structured LLM-Generated Feedback on Programming Assignment Performance
by: Mihaylova, Tsvetomila, et al.
Published: (2026)
by: Mihaylova, Tsvetomila, et al.
Published: (2026)
AI-Generated Slides: Are They Good? Can Students Tell?
by: Leinonen, Juho, et al.
Published: (2026)
by: Leinonen, Juho, et al.
Published: (2026)
Experiences from Integrating Large Language Model Chatbots into the Classroom
by: Hellas, Arto, et al.
Published: (2024)
by: Hellas, Arto, et al.
Published: (2024)
Teaching Language Models How to Code Like Learners: Conversational Serialization for Student Simulation
by: Koutcheme, Charles, et al.
Published: (2026)
by: Koutcheme, Charles, et al.
Published: (2026)
On the Opportunities of Large Language Models for Programming Process Data
by: Edwards, John, et al.
Published: (2024)
by: Edwards, John, et al.
Published: (2024)
LLM-itation is the Sincerest Form of Data: Generating Synthetic Buggy Code Submissions for Computing Education
by: Leinonen, Juho, et al.
Published: (2024)
by: Leinonen, Juho, et al.
Published: (2024)
61A Bot Report: AI Assistants in CS1 Save Students Homework Time and Reduce Demands on Staff. (Now What?)
by: Zamfirescu-Pereira, J. D., et al.
Published: (2024)
by: Zamfirescu-Pereira, J. D., et al.
Published: (2024)
Evaluating Contextually Personalized Programming Exercises Created with Generative AI
by: Logacheva, Evanfiya, et al.
Published: (2024)
by: Logacheva, Evanfiya, et al.
Published: (2024)
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge
by: Koutcheme, Charles, et al.
Published: (2024)
by: Koutcheme, Charles, et al.
Published: (2024)
Benchmarking Educational Program Repair
by: Koutcheme, Charles, et al.
Published: (2024)
by: Koutcheme, Charles, et al.
Published: (2024)
Can We Improve Educational Diagram Generation with In-Context Examples? Not if a Hallucination Spoils the Bunch
by: Logacheva, Evanfiya, et al.
Published: (2026)
by: Logacheva, Evanfiya, et al.
Published: (2026)
Pensieve Discuss: Scalable Small-Group CS Tutoring System with AI
by: Yang, Yoonseok, et al.
Published: (2024)
by: Yang, Yoonseok, et al.
Published: (2024)
A Knowledge-Component-Based Methodology for Evaluating AI Assistants
by: Qi, Laryn, et al.
Published: (2024)
by: Qi, Laryn, et al.
Published: (2024)
Let's Ask AI About Their Programs: Exploring ChatGPT's Answers To Program Comprehension Questions
by: Lehtinen, Teemu, et al.
Published: (2024)
by: Lehtinen, Teemu, et al.
Published: (2024)
Synthetic Students: A Comparative Study of Bug Distribution Between Large Language Models and Computing Students
by: MacNeil, Stephen, et al.
Published: (2024)
by: MacNeil, Stephen, et al.
Published: (2024)
ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle
by: Miroyan, Mihran, et al.
Published: (2025)
by: Miroyan, Mihran, et al.
Published: (2025)
Evaluating Language Models for Generating and Judging Programming Feedback
by: Koutcheme, Charles, et al.
Published: (2024)
by: Koutcheme, Charles, et al.
Published: (2024)
Using Large Language Models to Enhance Programming Error Messages
by: Leinonen, Juho, et al.
Published: (2022)
by: Leinonen, Juho, et al.
Published: (2022)
Comparing Code Explanations Created by Students and Large Language Models
by: Leinonen, Juho, et al.
Published: (2023)
by: Leinonen, Juho, et al.
Published: (2023)
"Like a Nesting Doll": Analyzing Recursion Analogies Generated by CS Students using Large Language Models
by: Bernstein, Seth, et al.
Published: (2024)
by: Bernstein, Seth, et al.
Published: (2024)
LeanTutor: Towards a Verified AI Mathematical Proof Tutor
by: Patel, Manooshree, et al.
Published: (2025)
by: Patel, Manooshree, et al.
Published: (2025)
Cross-lingual Human-Preference Alignment for Neural Machine Translation with Direct Quality Optimization
by: Uhlig, Kaden, et al.
Published: (2024)
by: Uhlig, Kaden, et al.
Published: (2024)
Comparing the Utility, Preference, and Performance of Course Material Search Functionality and Retrieval-Augmented Generation Large Language Model (RAG-LLM) AI Chatbots in Information-Seeking Tasks
by: Pasquarelli, Leonardo, et al.
Published: (2024)
by: Pasquarelli, Leonardo, et al.
Published: (2024)
Overview of Web Application Performance Optimization Techniques
by: Vepsäläinen, Juho, et al.
Published: (2024)
by: Vepsäläinen, Juho, et al.
Published: (2024)
Facilitating Instructors-LLM Collaboration for Problem Design in Introductory Programming Classrooms
by: Hoq, Muntasir, et al.
Published: (2025)
by: Hoq, Muntasir, et al.
Published: (2025)
"Sometimes You Just Gotta Risk It for the Biscuit": A Portrait of Student Risk-Taking
by: Leinonen, Juho, et al.
Published: (2024)
by: Leinonen, Juho, et al.
Published: (2024)
Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
by: Stamatis, Caitlin A., et al.
Published: (2026)
by: Stamatis, Caitlin A., et al.
Published: (2026)
How Large Language Models Are Changing MOOC Essay Answers: A Comparison of Pre- and Post-LLM Responses
by: Leppänen, Leo, et al.
Published: (2025)
by: Leppänen, Leo, et al.
Published: (2025)
NEFMind: Parameter-Efficient Fine-Tuning of Open-Source LLMs for Telecom APIs Automation
by: Khan, Zainab, et al.
Published: (2025)
by: Khan, Zainab, et al.
Published: (2025)
Automated Knowledge Component Generation for Interpretable Knowledge Tracing in Coding Problems
by: Duan, Zhangqi, et al.
Published: (2025)
by: Duan, Zhangqi, et al.
Published: (2025)
Probing the Unknown: Exploring Student Interactions with Probeable Problems at Scale in Introductory Programming
by: Denny, Paul, et al.
Published: (2025)
by: Denny, Paul, et al.
Published: (2025)
Mapping the Exploitation Surface: A 10,000-Trial Taxonomy of What Makes LLM Agents Exploit Vulnerabilities
by: Mouzouni, Charafeddine
Published: (2026)
by: Mouzouni, Charafeddine
Published: (2026)
Quantitative Analysis of AI-Generated Texts in Academic Research: A Study of AI Presence in Arxiv Submissions using AI Detection Tool
by: Akram, Arslan
Published: (2024)
by: Akram, Arslan
Published: (2024)
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
by: Bao, Qiming, et al.
Published: (2026)
by: Bao, Qiming, et al.
Published: (2026)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024)
by: Norouzi, Narges, et al.
Published: (2024)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
by: Cavagnero, Niccolò, et al.
Published: (2026)
by: Cavagnero, Niccolò, et al.
Published: (2026)
Narrowing the Gap: Supervised Fine-Tuning of Open-Source LLMs as a Viable Alternative to Proprietary Models for Pedagogical Tools
by: Solano, Lorenzo Lee, et al.
Published: (2025)
by: Solano, Lorenzo Lee, et al.
Published: (2025)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
by: Skorobogat, Ronald, et al.
Published: (2026)
by: Skorobogat, Ronald, et al.
Published: (2026)
Similar Items
-
Personalized Worked Example Generation from Student Code Submissions Using Pattern-based Knowledge Components
by: Pitts, Griffin, et al.
Published: (2026) -
Retrieval-Augmented Tutoring for Algorithm Tracing and Problem-Solving in AI Education
by: Jain, Mragisha, et al.
Published: (2026) -
The Effects of Structured LLM-Generated Feedback on Programming Assignment Performance
by: Mihaylova, Tsvetomila, et al.
Published: (2026) -
AI-Generated Slides: Are They Good? Can Students Tell?
by: Leinonen, Juho, et al.
Published: (2026) -
Experiences from Integrating Large Language Model Chatbots into the Classroom
by: Hellas, Arto, et al.
Published: (2024)