Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
Fuente:
arXiv
Saved in:
| Main Authors: | Thomas, Danielle R., Borchers, Conrad, Vanacore, Kirk P., Koedinger, Kenneth R., Kizilcec, René F. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Agreement: Rethinking Ground Truth in Educational AI Annotation
by: Thomas, Danielle R., et al.
Published: (2025)
by: Thomas, Danielle R., et al.
Published: (2025)
Who Decides in AI-Mediated Learning? The Agency Allocation Framework
by: Borchers, Conrad, et al.
Published: (2026)
by: Borchers, Conrad, et al.
Published: (2026)
How well do Large Language Models Recognize Instructional Moves? Establishing Baselines for Foundation Models in Educational Discourse
by: Vanacore, Kirk, et al.
Published: (2025)
by: Vanacore, Kirk, et al.
Published: (2025)
Augmenting Human-Annotated Training Data with Large Language Model Generation and Distillation in Open-Response Assessment
by: Borchers, Conrad, et al.
Published: (2025)
by: Borchers, Conrad, et al.
Published: (2025)
Brief but Impactful: How Human Tutoring Interactions Shape Engagement in Online Learning
by: Borchers, Conrad, et al.
Published: (2026)
by: Borchers, Conrad, et al.
Published: (2026)
The Digital Divide in Generative AI: Evidence from Large Language Model Use in College Admissions Essays
by: Lee, Jinsook, et al.
Published: (2026)
by: Lee, Jinsook, et al.
Published: (2026)
LLM-Generated Feedback Supports Learning If Learners Choose to Use It
by: Thomas, Danielle R., et al.
Published: (2025)
by: Thomas, Danielle R., et al.
Published: (2025)
Starting Seatwork Earlier as a Valid Measure of Student Engagement
by: Gurung, Ashish, et al.
Published: (2025)
by: Gurung, Ashish, et al.
Published: (2025)
Million Tutoring Moves (MTM): An Open Multimodal Dataset for the Science of Tutoring
by: Kizilcec, René, et al.
Published: (2026)
by: Kizilcec, René, et al.
Published: (2026)
Improving Assessment of Tutoring Practices using Retrieval-Augmented Generation
by: Han, Zifei FeiFei, et al.
Published: (2024)
by: Han, Zifei FeiFei, et al.
Published: (2024)
Optimizing LLM Annotation of Classroom Discourse through Multi-Agent Orchestration
by: Ahtisham, Bakhtawar, et al.
Published: (2026)
by: Ahtisham, Bakhtawar, et al.
Published: (2026)
"Don't Forget the Teachers": Towards an Educator-Centered Understanding of Harms from Large Language Models in Education
by: Harvey, Emma, et al.
Published: (2025)
by: Harvey, Emma, et al.
Published: (2025)
Leveraging LLMs to Assess Tutor Moves in Real-Life Dialogues: A Feasibility Study
by: Thomas, Danielle R., et al.
Published: (2025)
by: Thomas, Danielle R., et al.
Published: (2025)
Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training
by: Kakarla, Sanjit, et al.
Published: (2025)
by: Kakarla, Sanjit, et al.
Published: (2025)
Does the TalkMoves Codebook Generalize to One-on-One Tutoring and Multimodal Interaction?
by: Focsan, Corina Luca, et al.
Published: (2026)
by: Focsan, Corina Luca, et al.
Published: (2026)
Improving Hybrid Human-AI Tutoring by Differentiating Human Tutor Roles Based on Student Needs
by: Gurung, Ashish, et al.
Published: (2026)
by: Gurung, Ashish, et al.
Published: (2026)
AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics
by: Ahtisham, Bakhtawar, et al.
Published: (2025)
by: Ahtisham, Bakhtawar, et al.
Published: (2025)
Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale
by: Hedley, Daryl, et al.
Published: (2026)
by: Hedley, Daryl, et al.
Published: (2026)
From Heuristics to Analytics: Forecasting Effort and Progress in Online Learning
by: Qiu, Eric S., et al.
Published: (2026)
by: Qiu, Eric S., et al.
Published: (2026)
When the Past Misleads: Rethinking Training Data Expansion Under Temporal Distribution Shifts
by: Yao, Chengyuan, et al.
Published: (2025)
by: Yao, Chengyuan, et al.
Published: (2025)
Coasting Through Class: Learning Opportunity Loss from Practice Avoidance During Individual Seatwork
by: Gurung, Ashish, et al.
Published: (2026)
by: Gurung, Ashish, et al.
Published: (2026)
Using Large Language Models to Assess Tutors' Performance in Reacting to Students Making Math Errors
by: Kakarla, Sanjit, et al.
Published: (2024)
by: Kakarla, Sanjit, et al.
Published: (2024)
Synthetic Data and the Shifting Ground of Truth
by: Offenhuber, Dietmar
Published: (2025)
by: Offenhuber, Dietmar
Published: (2025)
Toward Trait-Aware Learning Analytics
by: Borchers, Conrad, et al.
Published: (2026)
by: Borchers, Conrad, et al.
Published: (2026)
Fairness Hub Technical Briefs: Definition and Detection of Distribution Shift
by: Acevedo, Nicolas, et al.
Published: (2024)
by: Acevedo, Nicolas, et al.
Published: (2024)
A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
by: Harvey, Emma, et al.
Published: (2025)
by: Harvey, Emma, et al.
Published: (2025)
Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts
by: Lee, Jinsook, et al.
Published: (2026)
by: Lee, Jinsook, et al.
Published: (2026)
LLM Reasoning Predicts When Models Are Right: Evidence from Coding Classroom Discourse
by: Ahtisham, Bakhtawar, et al.
Published: (2026)
by: Ahtisham, Bakhtawar, et al.
Published: (2026)
Do Tutors Learn from Equity Training and Can Generative AI Assess It?
by: Thomas, Danielle R., et al.
Published: (2024)
by: Thomas, Danielle R., et al.
Published: (2024)
The Life Cycle of Large Language Models: A Review of Biases in Education
by: Lee, Jinsook, et al.
Published: (2024)
by: Lee, Jinsook, et al.
Published: (2024)
Does Multiple Choice Have a Future in the Age of Generative AI? A Posttest-only RCT
by: Thomas, Danielle R., et al.
Published: (2024)
by: Thomas, Danielle R., et al.
Published: (2024)
An Integrated Platform for Studying Learning with Intelligent Tutoring Systems: CTAT+TutorShop
by: Aleven, Vincent, et al.
Published: (2025)
by: Aleven, Vincent, et al.
Published: (2025)
Reflection on Purpose Changes Students' Academic Interests: A Scalable Intervention in an Online Course Catalog
by: Chen, Youjie, et al.
Published: (2024)
by: Chen, Youjie, et al.
Published: (2024)
Teachers' Perceived Benefits and Risks of AI Across Fifty-Five Countries: An Audit of LLM Alignment and Steerability
by: Tao, Yan, et al.
Published: (2026)
by: Tao, Yan, et al.
Published: (2026)
Tutor Move Taxonomy: A Theory-Aligned Framework for Analyzing Instructional Moves in Tutoring
by: Zhou, Zhuqian, et al.
Published: (2026)
by: Zhou, Zhuqian, et al.
Published: (2026)
Representation Learning to Study Temporal Dynamics in Tutorial Scaffolding
by: Borchers, Conrad, et al.
Published: (2026)
by: Borchers, Conrad, et al.
Published: (2026)
How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures
by: Zhang, Shan, et al.
Published: (2026)
by: Zhang, Shan, et al.
Published: (2026)
Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation
by: Lee, Jinsook, et al.
Published: (2026)
by: Lee, Jinsook, et al.
Published: (2026)
Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident
by: Borchers, Conrad, et al.
Published: (2026)
by: Borchers, Conrad, et al.
Published: (2026)
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
by: Du, Yishan, et al.
Published: (2025)
by: Du, Yishan, et al.
Published: (2025)
Similar Items
-
Beyond Agreement: Rethinking Ground Truth in Educational AI Annotation
by: Thomas, Danielle R., et al.
Published: (2025) -
Who Decides in AI-Mediated Learning? The Agency Allocation Framework
by: Borchers, Conrad, et al.
Published: (2026) -
How well do Large Language Models Recognize Instructional Moves? Establishing Baselines for Foundation Models in Educational Discourse
by: Vanacore, Kirk, et al.
Published: (2025) -
Augmenting Human-Annotated Training Data with Large Language Model Generation and Distillation in Open-Response Assessment
by: Borchers, Conrad, et al.
Published: (2025) -
Brief but Impactful: How Human Tutoring Interactions Shape Engagement in Online Learning
by: Borchers, Conrad, et al.
Published: (2026)