From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Weikang, Ren, Houxing, Pan, Junting, Zhou, Aojun, Wang, Ke, Lu, Zimu, Yang, Yunqiao, Hu, Yuxuan, Wei, Linda, Zhan, Mingjie, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
by: Yang, Yunqiao, et al.
Published: (2025)
by: Yang, Yunqiao, et al.
Published: (2025)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
Alignment with Fill-In-the-Middle for Enhancing Code Generation
by: Ren, Houxing, et al.
Published: (2025)
by: Ren, Houxing, et al.
Published: (2025)
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
Edit-Based Refinement for Parallel Masked Diffusion Language Models
by: Ren, Houxing, et al.
Published: (2026)
by: Ren, Houxing, et al.
Published: (2026)
SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics
by: Yang, Yunqiao, et al.
Published: (2026)
by: Yang, Yunqiao, et al.
Published: (2026)
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
by: Ren, Houxing, et al.
Published: (2024)
by: Ren, Houxing, et al.
Published: (2024)
Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format Reasoning
by: Ren, Houxing, et al.
Published: (2026)
by: Ren, Houxing, et al.
Published: (2026)
FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation
by: Lu, Zimu, et al.
Published: (2026)
by: Lu, Zimu, et al.
Published: (2026)
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
by: Wang, Ke, et al.
Published: (2024)
by: Wang, Ke, et al.
Published: (2024)
VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
Empowering Character-level Text Infilling by Eliminating Sub-Tokens
by: Ren, Houxing, et al.
Published: (2024)
by: Ren, Houxing, et al.
Published: (2024)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
by: Shi, Weikang, et al.
Published: (2025)
by: Shi, Weikang, et al.
Published: (2025)
LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
by: Macina, Jakub, et al.
Published: (2025)
by: Macina, Jakub, et al.
Published: (2025)
LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring
by: Lee, Unggi, et al.
Published: (2026)
by: Lee, Unggi, et al.
Published: (2026)
NODI: Out-Of-Distribution Detection with Noise from Diffusion
by: Zhou, Jingqiu, et al.
Published: (2024)
by: Zhou, Jingqiu, et al.
Published: (2024)
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
Genetic Auto-prompt Learning for Pre-trained Code Intelligence Language Models
by: Feng, Chengzhe, et al.
Published: (2024)
by: Feng, Chengzhe, et al.
Published: (2024)
SpiritSight Agent: Advanced GUI Agent with One Look
by: Huang, Zhiyuan, et al.
Published: (2025)
by: Huang, Zhiyuan, et al.
Published: (2025)
Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions
by: Cheng, Ziming, et al.
Published: (2025)
by: Cheng, Ziming, et al.
Published: (2025)
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
by: Hazra, Rima, et al.
Published: (2026)
by: Hazra, Rima, et al.
Published: (2026)
From Untamed Black Box to Interpretable Pedagogical Orchestration: The Ensemble of Specialized LLMs Architecture for Adaptive Tutoring
by: Kadir, Nizam
Published: (2026)
by: Kadir, Nizam
Published: (2026)
Hidden temperature in the KMP model
by: De Masi, Anna, et al.
Published: (2023)
by: De Masi, Anna, et al.
Published: (2023)
Convergence of the KMP model to the KPZ equation
by: Barraquand, Guillaume, et al.
Published: (2025)
by: Barraquand, Guillaume, et al.
Published: (2025)
Bridging VLM and KMP: Enabling Fine-grained robotic manipulation via Semantic Keypoints Representation
by: Zhu, Junjie, et al.
Published: (2025)
by: Zhu, Junjie, et al.
Published: (2025)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025)
by: Jiang, Yilin, et al.
Published: (2025)
MolViBench: Evaluating LLMs on Molecular Vibe Coding
by: Li, Jiatong, et al.
Published: (2026)
by: Li, Jiatong, et al.
Published: (2026)
BIPED: Pedagogically Informed Tutoring System for ESL Education
by: Kwon, Soonwoo, et al.
Published: (2024)
by: Kwon, Soonwoo, et al.
Published: (2024)
PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning
by: Chang, Qikai, et al.
Published: (2026)
by: Chang, Qikai, et al.
Published: (2026)
The Impact of Artificial Intelligence Literacy on Doctoral Students' Innovative Behaviour From the Perspective of Technology Affordance
by: Lu Weikang, et al.
Published: (2025)
by: Lu Weikang, et al.
Published: (2025)
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
CourseAssist: Pedagogically Appropriate AI Tutor for Computer Science Education
by: Feng, Ty, et al.
Published: (2024)
by: Feng, Ty, et al.
Published: (2024)
Intertwining and propagation of mixtures for generalized KMP models and harmonic models
by: Giardinà, Cristian, et al.
Published: (2024)
by: Giardinà, Cristian, et al.
Published: (2024)
Practical KMP/BM Style Pattern-Matching on Indeterminate Strings
by: Dehghani, Hossein, et al.
Published: (2022)
by: Dehghani, Hossein, et al.
Published: (2022)
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
by: Srinivasa, Rakshith S, et al.
Published: (2025)
by: Srinivasa, Rakshith S, et al.
Published: (2025)
Similar Items
-
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
by: Lu, Zimu, et al.
Published: (2025) -
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
by: Yang, Yunqiao, et al.
Published: (2025) -
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
by: Lu, Zimu, et al.
Published: (2024) -
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning
by: Lu, Zimu, et al.
Published: (2024) -
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
by: Lu, Zimu, et al.
Published: (2024)