Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
Fuente:
arXiv
Saved in:
| Main Authors: | Kochmar, Ekaterina, Maurya, Kaushal Kumar, Petukhova, Kseniia, Srivatsa, KV Aditya, Tack, Anaïs, Vasselli, Justin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
Intent Matters: Enhancing AI Tutoring with Fine-Grained Pedagogical Intent Annotation
by: Petukhova, Kseniia, et al.
Published: (2025)
by: Petukhova, Kseniia, et al.
Published: (2025)
AITutor-EvalKit: Exploring the Capabilities of AI Tutors
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Can LLMs Reliably Simulate Real Students' Abilities in Mathematics and Reading Comprehension?
by: Srivatsa, KV Aditya, et al.
Published: (2025)
by: Srivatsa, KV Aditya, et al.
Published: (2025)
Towards Reward Modeling for AI Tutors in Math Mistake Remediation
by: Petukhova, Kseniia, et al.
Published: (2026)
by: Petukhova, Kseniia, et al.
Published: (2026)
SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing
by: Srivatsa, KV Aditya, et al.
Published: (2024)
by: Srivatsa, KV Aditya, et al.
Published: (2024)
LLMs cannot spot math errors, even when allowed to peek into the solution
by: Srivatsa, KV Aditya, et al.
Published: (2025)
by: Srivatsa, KV Aditya, et al.
Published: (2025)
Pedagogy-driven Evaluation of Generative AI-powered Intelligent Tutoring Systems
by: Maurya, Kaushal Kumar, et al.
Published: (2025)
by: Maurya, Kaushal Kumar, et al.
Published: (2025)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
by: Tonga, Junior Cedric, et al.
Published: (2025)
by: Tonga, Junior Cedric, et al.
Published: (2025)
What Makes Math Word Problems Challenging for LLMs?
by: Srivatsa, KV Aditya, et al.
Published: (2024)
by: Srivatsa, KV Aditya, et al.
Published: (2024)
A Fully Automated Pipeline for Conversational Discourse Annotation: Tree Scheme Generation and Labeling with Large Language Models
by: Petukhova, Kseniia, et al.
Published: (2025)
by: Petukhova, Kseniia, et al.
Published: (2025)
PetKaz at SemEval-2024 Task 8: Can Linguistics Capture the Specifics of LLM-generated Text?
by: Petukhova, Kseniia, et al.
Published: (2024)
by: Petukhova, Kseniia, et al.
Published: (2024)
PetKaz at SemEval-2024 Task 3: Advancing Emotion Classification with an LLM for Emotion-Cause Pair Extraction in Conversations
by: Kazakov, Roman, et al.
Published: (2024)
by: Kazakov, Roman, et al.
Published: (2024)
BD at BEA 2025 Shared Task: MPNet Ensembles for Pedagogical Mistake Identification and Localization in AI Tutor Responses
by: Rohan, Shadman, et al.
Published: (2025)
by: Rohan, Shadman, et al.
Published: (2025)
RETUYT-INCO at BEA 2025 Shared Task: How Far Can Lightweight Models Go in AI-powered Tutor Evaluation?
by: Góngora, Santiago, et al.
Published: (2025)
by: Góngora, Santiago, et al.
Published: (2025)
NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Opportunities and Challenges of LLMs in Education: An NLP Perspective
by: Vajjala, Sowmya, et al.
Published: (2025)
by: Vajjala, Sowmya, et al.
Published: (2025)
LLMs in Education: Novel Perspectives, Challenges, and Opportunities
by: Alhafni, Bashar, et al.
Published: (2024)
by: Alhafni, Bashar, et al.
Published: (2024)
MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors
by: Hikal, Baraa, et al.
Published: (2025)
by: Hikal, Baraa, et al.
Published: (2025)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
by: Liang, Junhong, et al.
Published: (2026)
by: Liang, Junhong, et al.
Published: (2026)
The ADAIO System at the BEA-2023 Shared Task on Generating AI Teacher Responses in Educational Dialogues
by: Adigwe, Adaeze, et al.
Published: (2023)
by: Adigwe, Adaeze, et al.
Published: (2023)
UNVEILING: What Makes Linguistics Olympiad Puzzles Tricky for LLMs?
by: Choudhary, Mukund, et al.
Published: (2025)
by: Choudhary, Mukund, et al.
Published: (2025)
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
by: Hazra, Rima, et al.
Published: (2026)
by: Hazra, Rima, et al.
Published: (2026)
REFeREE: A REference-FREE Model-Based Metric for Text Simplification
by: Huang, Yichen, et al.
Published: (2024)
by: Huang, Yichen, et al.
Published: (2024)
Teaching Through Analogies: A Modular Pipeline for Educational Analogy Generation
by: Barakat, Mariam, et al.
Published: (2026)
by: Barakat, Mariam, et al.
Published: (2026)
Sakura at BEA 2026 Shared Task 1: What Makes Vocabulary Difficult?
by: Nohejl, Adam, et al.
Published: (2026)
by: Nohejl, Adam, et al.
Published: (2026)
Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error Correction
by: Goto, Takumi, et al.
Published: (2024)
by: Goto, Takumi, et al.
Published: (2024)
Measuring the Robustness of Reference-Free Dialogue Evaluation Systems
by: Vasselli, Justin, et al.
Published: (2025)
by: Vasselli, Justin, et al.
Published: (2025)
CourseAssist: Pedagogically Appropriate AI Tutor for Computer Science Education
by: Feng, Ty, et al.
Published: (2024)
by: Feng, Ty, et al.
Published: (2024)
What Makes Cryptic Crosswords Challenging for LLMs?
by: Sadallah, Abdelrahman, et al.
Published: (2024)
by: Sadallah, Abdelrahman, et al.
Published: (2024)
RETUYT-INCO at BEA 2026 Shared Task 2: Meta-prompting in Rubric-based Scoring for German
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
Are LLMs Good Cryptic Crossword Solvers?
by: Sadallah, Abdelrahman, et al.
Published: (2024)
by: Sadallah, Abdelrahman, et al.
Published: (2024)
An Experience Report on a Pedagogically Controlled, Curriculum-Constrained AI Tutor for SE Education
by: Happe, Lucia, et al.
Published: (2025)
by: Happe, Lucia, et al.
Published: (2025)
Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation
by: Zhang, Luke, et al.
Published: (2026)
by: Zhang, Luke, et al.
Published: (2026)
LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring
by: Lee, Unggi, et al.
Published: (2026)
by: Lee, Unggi, et al.
Published: (2026)
BIPED: Pedagogically Informed Tutoring System for ESL Education
by: Kwon, Soonwoo, et al.
Published: (2024)
by: Kwon, Soonwoo, et al.
Published: (2024)
Multilingual Dialogue Generation and Localization with Dialogue Act Scripting
by: Vasselli, Justin, et al.
Published: (2025)
by: Vasselli, Justin, et al.
Published: (2025)
mucAI at BAREC Shared Task 2025: Towards Uncertainty Aware Arabic Readability Assessment
by: Abdou, Ahmed
Published: (2025)
by: Abdou, Ahmed
Published: (2025)
PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning
by: Chang, Qikai, et al.
Published: (2026)
by: Chang, Qikai, et al.
Published: (2026)
Similar Items
-
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
by: Maurya, Kaushal Kumar, et al.
Published: (2024) -
Intent Matters: Enhancing AI Tutoring with Fine-Grained Pedagogical Intent Annotation
by: Petukhova, Kseniia, et al.
Published: (2025) -
AITutor-EvalKit: Exploring the Capabilities of AI Tutors
by: Naeem, Numaan, et al.
Published: (2025) -
Can LLMs Reliably Simulate Real Students' Abilities in Mathematics and Reading Comprehension?
by: Srivatsa, KV Aditya, et al.
Published: (2025) -
Towards Reward Modeling for AI Tutors in Math Mistake Remediation
by: Petukhova, Kseniia, et al.
Published: (2026)