MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
Fuente:
arXiv
Guardado en:
| Autores principales: | Macina, Jakub, Daheim, Nico, Hakimi, Ido, Kapur, Manu, Gurevych, Iryna, Sachan, Mrinmaya |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
por: Daheim, Nico, et al.
Publicado: (2024)
por: Daheim, Nico, et al.
Publicado: (2024)
From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning
por: Dinucu-Jianu, David, et al.
Publicado: (2025)
por: Dinucu-Jianu, David, et al.
Publicado: (2025)
Towards the Pedagogical Steering of Large Language Models for Tutoring: A Case Study with Modeling Productive Failure
por: Puech, Romain, et al.
Publicado: (2024)
por: Puech, Romain, et al.
Publicado: (2024)
Book2Dial: Generating Teacher-Student Interactions from Textbooks for Cost-Effective Development of Educational Chatbots
por: Wang, Junling, et al.
Publicado: (2024)
por: Wang, Junling, et al.
Publicado: (2024)
Uncertainty-Aware Decoding with Minimum Bayes Risk
por: Daheim, Nico, et al.
Publicado: (2025)
por: Daheim, Nico, et al.
Publicado: (2025)
Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models
por: Wang, Junling, et al.
Publicado: (2025)
por: Wang, Junling, et al.
Publicado: (2025)
Token Weighting for Long-Range Language Modeling
por: Helm, Falko, et al.
Publicado: (2025)
por: Helm, Falko, et al.
Publicado: (2025)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
por: Chowdhury, Sankalan Pal, et al.
Publicado: (2024)
por: Chowdhury, Sankalan Pal, et al.
Publicado: (2024)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
por: Tonga, Junior Cedric, et al.
Publicado: (2025)
por: Tonga, Junior Cedric, et al.
Publicado: (2025)
PATS: Personality-Aware Teaching Strategies with Large Language Model Tutors
por: Rooein, Donya, et al.
Publicado: (2026)
por: Rooein, Donya, et al.
Publicado: (2026)
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
por: Hazra, Rima, et al.
Publicado: (2026)
por: Hazra, Rima, et al.
Publicado: (2026)
LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring
por: Lee, Unggi, et al.
Publicado: (2026)
por: Lee, Unggi, et al.
Publicado: (2026)
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
por: Yang, Tengchao, et al.
Publicado: (2025)
por: Yang, Tengchao, et al.
Publicado: (2025)
Socratic Reasoning Improves Positive Text Rewriting
por: Goel, Anmol, et al.
Publicado: (2024)
por: Goel, Anmol, et al.
Publicado: (2024)
Model Merging by Uncertainty-Based Gradient Matching
por: Daheim, Nico, et al.
Publicado: (2023)
por: Daheim, Nico, et al.
Publicado: (2023)
BIPED: Pedagogically Informed Tutoring System for ESL Education
por: Kwon, Soonwoo, et al.
Publicado: (2024)
por: Kwon, Soonwoo, et al.
Publicado: (2024)
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
por: Srinivasa, Rakshith S, et al.
Publicado: (2025)
por: Srinivasa, Rakshith S, et al.
Publicado: (2025)
MathBuddy: A Multimodal System for Affective Math Tutoring
por: Kar, Debanjana, et al.
Publicado: (2025)
por: Kar, Debanjana, et al.
Publicado: (2025)
Can Vision-Language Models Solve Visual Math Equations?
por: Choudhury, Monjoy Narayan, et al.
Publicado: (2025)
por: Choudhury, Monjoy Narayan, et al.
Publicado: (2025)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
por: Opedal, Andreas, et al.
Publicado: (2024)
por: Opedal, Andreas, et al.
Publicado: (2024)
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
por: Maurya, Kaushal Kumar, et al.
Publicado: (2024)
por: Maurya, Kaushal Kumar, et al.
Publicado: (2024)
DeepTutor: Towards Agentic Personalized Tutoring
por: Zhao, Bingxi, et al.
Publicado: (2026)
por: Zhao, Bingxi, et al.
Publicado: (2026)
From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
por: Shi, Weikang, et al.
Publicado: (2026)
por: Shi, Weikang, et al.
Publicado: (2026)
A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
por: Shelmanov, Artem, et al.
Publicado: (2025)
por: Shelmanov, Artem, et al.
Publicado: (2025)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
por: Do, Heejin, et al.
Publicado: (2026)
por: Do, Heejin, et al.
Publicado: (2026)
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
por: Kochmar, Ekaterina, et al.
Publicado: (2025)
por: Kochmar, Ekaterina, et al.
Publicado: (2025)
Preemptive Detection and Correction of Misaligned Actions in LLM Agents
por: Fang, Haishuo, et al.
Publicado: (2024)
por: Fang, Haishuo, et al.
Publicado: (2024)
World Models for Math Story Problems
por: Opedal, Andreas, et al.
Publicado: (2023)
por: Opedal, Andreas, et al.
Publicado: (2023)
Probing for Arithmetic Errors in Language Models
por: Sun, Yucheng, et al.
Publicado: (2025)
por: Sun, Yucheng, et al.
Publicado: (2025)
Heterogeneous Debate Engine: Identity-Grounded Cognitive Architecture for Resilient LLM-Based Ethical Tutoring
por: Masłowski, Jakub, et al.
Publicado: (2026)
por: Masłowski, Jakub, et al.
Publicado: (2026)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
por: Baumgärtner, Tim, et al.
Publicado: (2026)
por: Baumgärtner, Tim, et al.
Publicado: (2026)
Large Language Models Approach Expert Pedagogical Quality in Math Tutoring but Differ in Instructional and Linguistic Profiles
por: Abdulsalam, Ramatu Oiza, et al.
Publicado: (2025)
por: Abdulsalam, Ramatu Oiza, et al.
Publicado: (2025)
Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors
por: Wang, Jian, et al.
Publicado: (2025)
por: Wang, Jian, et al.
Publicado: (2025)
Towards Actionable Pedagogical Feedback: A Multi-Perspective Analysis of Mathematics Teaching and Tutoring Dialogue
por: Naim, Jannatun, et al.
Publicado: (2025)
por: Naim, Jannatun, et al.
Publicado: (2025)
How to Weight Multitask Finetuning? Fast Previews via Bayesian Model-Merging
por: Maldonado, Hugo Monzón, et al.
Publicado: (2024)
por: Maldonado, Hugo Monzón, et al.
Publicado: (2024)
Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset
por: Zhou, Zhuqian, et al.
Publicado: (2026)
por: Zhou, Zhuqian, et al.
Publicado: (2026)
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
por: Rizvi, Md Imbesat Hassan, et al.
Publicado: (2024)
por: Rizvi, Md Imbesat Hassan, et al.
Publicado: (2024)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
por: Paul, Indraneil, et al.
Publicado: (2024)
por: Paul, Indraneil, et al.
Publicado: (2024)
How to Engage Your Readers? Generating Guiding Questions to Promote Active Reading
por: Cui, Peng, et al.
Publicado: (2024)
por: Cui, Peng, et al.
Publicado: (2024)
GPT-4 as a Homework Tutor can Improve Student Engagement and Learning Outcomes
por: Vanzo, Alessandro, et al.
Publicado: (2024)
por: Vanzo, Alessandro, et al.
Publicado: (2024)
Ejemplares similares
-
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
por: Daheim, Nico, et al.
Publicado: (2024) -
From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning
por: Dinucu-Jianu, David, et al.
Publicado: (2025) -
Towards the Pedagogical Steering of Large Language Models for Tutoring: A Case Study with Modeling Productive Failure
por: Puech, Romain, et al.
Publicado: (2024) -
Book2Dial: Generating Teacher-Student Interactions from Textbooks for Cost-Effective Development of Educational Chatbots
por: Wang, Junling, et al.
Publicado: (2024) -
Uncertainty-Aware Decoding with Minimum Bayes Risk
por: Daheim, Nico, et al.
Publicado: (2025)