LLM Reasoning Predicts When Models Are Right: Evidence from Coding Classroom Discourse
Fuente:
arXiv
Saved in:
| Main Authors: | Ahtisham, Bakhtawar, Vanacore, Kirk, Zhou, Zhuqian, Lee, Jinsook, Kizilcec, Rene F. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts
by: Lee, Jinsook, et al.
Published: (2026)
by: Lee, Jinsook, et al.
Published: (2026)
Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation
by: Lee, Jinsook, et al.
Published: (2026)
by: Lee, Jinsook, et al.
Published: (2026)
Optimizing LLM Annotation of Classroom Discourse through Multi-Agent Orchestration
by: Ahtisham, Bakhtawar, et al.
Published: (2026)
by: Ahtisham, Bakhtawar, et al.
Published: (2026)
AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics
by: Ahtisham, Bakhtawar, et al.
Published: (2025)
by: Ahtisham, Bakhtawar, et al.
Published: (2025)
Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset
by: Zhou, Zhuqian, et al.
Published: (2026)
by: Zhou, Zhuqian, et al.
Published: (2026)
How well do Large Language Models Recognize Instructional Moves? Establishing Baselines for Foundation Models in Educational Discourse
by: Vanacore, Kirk, et al.
Published: (2025)
by: Vanacore, Kirk, et al.
Published: (2025)
Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale
by: Hedley, Daryl, et al.
Published: (2026)
by: Hedley, Daryl, et al.
Published: (2026)
Tutor Move Taxonomy: A Theory-Aligned Framework for Analyzing Instructional Moves in Tutoring
by: Zhou, Zhuqian, et al.
Published: (2026)
by: Zhou, Zhuqian, et al.
Published: (2026)
Poor Alignment and Steerability of Large Language Models: Evidence from College Admission Essays
by: Lee, Jinsook, et al.
Published: (2025)
by: Lee, Jinsook, et al.
Published: (2025)
Million Tutoring Moves (MTM): An Open Multimodal Dataset for the Science of Tutoring
by: Kizilcec, René, et al.
Published: (2026)
by: Kizilcec, René, et al.
Published: (2026)
The Life Cycle of Large Language Models: A Review of Biases in Education
by: Lee, Jinsook, et al.
Published: (2024)
by: Lee, Jinsook, et al.
Published: (2024)
The Digital Divide in Generative AI: Evidence from Large Language Model Use in College Admissions Essays
by: Lee, Jinsook, et al.
Published: (2026)
by: Lee, Jinsook, et al.
Published: (2026)
Algorithms for College Admissions Decision Support: Impacts of Policy Change and Inherent Variability
by: Lee, Jinsook, et al.
Published: (2024)
by: Lee, Jinsook, et al.
Published: (2024)
Cultural Bias and Cultural Alignment of Large Language Models
by: Tao, Yan, et al.
Published: (2023)
by: Tao, Yan, et al.
Published: (2023)
Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
by: Thomas, Danielle R., et al.
Published: (2026)
by: Thomas, Danielle R., et al.
Published: (2026)
Shiksha Copilot: Teacher-AI Collaboration for Curating and Customizing Lesson Plans in Low-Resource Schools
by: Dennison, Deepak Varuvel, et al.
Published: (2025)
by: Dennison, Deepak Varuvel, et al.
Published: (2025)
Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reasoning-Component Classification
by: Noh, Jiho, et al.
Published: (2026)
by: Noh, Jiho, et al.
Published: (2026)
Thinking Out Loud: Do Reasoning Models Know When They're Right?
by: Zeng, Qingcheng, et al.
Published: (2025)
by: Zeng, Qingcheng, et al.
Published: (2025)
Do LLMs Signal When They're Right? Evidence from Neuron Agreement
by: Chen, Kang, et al.
Published: (2025)
by: Chen, Kang, et al.
Published: (2025)
The "Right" Discourse on Migration: Analysing Migration-Related Tweets in Right and Far-Right Political Movements
by: Chatterjee, Nishan, et al.
Published: (2025)
by: Chatterjee, Nishan, et al.
Published: (2025)
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
by: Wei, Jiaqi, et al.
Published: (2026)
by: Wei, Jiaqi, et al.
Published: (2026)
Code Execution as Grounded Supervision for LLM Reasoning
by: Jung, Dongwon, et al.
Published: (2025)
by: Jung, Dongwon, et al.
Published: (2025)
Does Algorithmic Uncertainty Sway Human Experts? Evidence from a Field Experiment in Selective College Admissions
by: Lee, Hansol, et al.
Published: (2026)
by: Lee, Hansol, et al.
Published: (2026)
The life cycle of large language models in education: A framework for understanding sources of bias
by: Jinsook Lee, et al.
Published: (2024)
by: Jinsook Lee, et al.
Published: (2024)
Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
by: Zhang, Anqi, et al.
Published: (2025)
by: Zhang, Anqi, et al.
Published: (2025)
When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs
by: Fekete, Marcell, et al.
Published: (2026)
by: Fekete, Marcell, et al.
Published: (2026)
Scientific Discourse Tagging for Evidence Extraction
by: Li, Xiangci, et al.
Published: (2019)
by: Li, Xiangci, et al.
Published: (2019)
When People are Floods: Analyzing Dehumanizing Metaphors in Immigration Discourse with Large Language Models
by: Mendelsohn, Julia, et al.
Published: (2025)
by: Mendelsohn, Julia, et al.
Published: (2025)
Assessing LLM Reasoning Through Implicit Causal Chain Discovery in Climate Discourse
by: Allein, Liesbeth, et al.
Published: (2025)
by: Allein, Liesbeth, et al.
Published: (2025)
When Do Language Models Endorse Limitations on Human Rights Principles?
by: Samway, Keenan, et al.
Published: (2026)
by: Samway, Keenan, et al.
Published: (2026)
Unraveling Misinformation Propagation in LLM Reasoning
by: Feng, Yiyang, et al.
Published: (2025)
by: Feng, Yiyang, et al.
Published: (2025)
Simulating Classroom Education with LLM-Empowered Agents
by: Zhang, Zheyuan, et al.
Published: (2024)
by: Zhang, Zheyuan, et al.
Published: (2024)
BeDiscovER: The Benchmark of Discourse Understanding in the Era of Reasoning Language Models
by: Li, Chuyuan, et al.
Published: (2025)
by: Li, Chuyuan, et al.
Published: (2025)
Towards an Analysis of Discourse and Interactional Pragmatic Reasoning Capabilities of Large Language Models
by: Robrecht, Amelie, et al.
Published: (2024)
by: Robrecht, Amelie, et al.
Published: (2024)
Modeling Topics and Sociolinguistic Variation in Code-Switched Discourse: Insights from Spanish-English and Spanish-Guaraní
by: Tyagi, Nemika, et al.
Published: (2025)
by: Tyagi, Nemika, et al.
Published: (2025)
LLM Agents Already Know When to Call Tools -- Even Without Reasoning
by: Sun, Chung-En, et al.
Published: (2026)
by: Sun, Chung-En, et al.
Published: (2026)
DIMSUM: Discourse in Mathematical Reasoning as a Supervision Module
by: Sharma, Krish, et al.
Published: (2025)
by: Sharma, Krish, et al.
Published: (2025)
The Realignment Problem: When Right becomes Wrong in LLMs
by: Sharma, Aakash Sen, et al.
Published: (2025)
by: Sharma, Aakash Sen, et al.
Published: (2025)
Learning When to Translate for Multilingual Reasoning
by: Kang, Deokhyung, et al.
Published: (2026)
by: Kang, Deokhyung, et al.
Published: (2026)
EviCare: Enhancing Diagnosis Prediction with Deep Model-Guided Evidence for In-Context Reasoning
by: Zhang, Hengyu, et al.
Published: (2026)
by: Zhang, Hengyu, et al.
Published: (2026)
Similar Items
-
Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts
by: Lee, Jinsook, et al.
Published: (2026) -
Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation
by: Lee, Jinsook, et al.
Published: (2026) -
Optimizing LLM Annotation of Classroom Discourse through Multi-Agent Orchestration
by: Ahtisham, Bakhtawar, et al.
Published: (2026) -
AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics
by: Ahtisham, Bakhtawar, et al.
Published: (2025) -
Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset
by: Zhou, Zhuqian, et al.
Published: (2026)