MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors
Fuente:
arXiv
Saved in:
| Main Authors: | Hikal, Baraa, Basem, Mohamed, Oshallah, Islam, Hamdi, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimized Quran Passage Retrieval Using an Expanded QA Dataset and Fine-Tuned Language Models
by: Basem, Mohamed, et al.
Published: (2024)
by: Basem, Mohamed, et al.
Published: (2024)
Few-Shot Prompting for Extractive Quranic QA with Instruction-Tuned LLMs
by: Basem, Mohamed, et al.
Published: (2025)
by: Basem, Mohamed, et al.
Published: (2025)
MSA at SemEval-2025 Task 3: High Quality Weak Labeling and LLM Ensemble Verification for Multilingual Hallucination Detection
by: Hikal, Baraa, et al.
Published: (2025)
by: Hikal, Baraa, et al.
Published: (2025)
Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction
by: Basem, Mohamed, et al.
Published: (2025)
by: Basem, Mohamed, et al.
Published: (2025)
Cross-Language Approach for Quranic QA
by: Oshallah, Islam, et al.
Published: (2025)
by: Oshallah, Islam, et al.
Published: (2025)
!MSA at BAREC Shared Task 2025: Ensembling Arabic Transformers for Readability Assessment
by: Basem, Mohamed, et al.
Published: (2025)
by: Basem, Mohamed, et al.
Published: (2025)
!MSA at AraHealthQA 2025 Shared Task: Enhancing LLM Performance for Arabic Clinical Question Answering through Prompt Engineering and Ensemble Learning
by: Tarek, Mohamed, et al.
Published: (2025)
by: Tarek, Mohamed, et al.
Published: (2025)
Few-Shot Optimized Framework for Hallucination Detection in Resource-Limited NLP Systems
by: Hikal, Baraa, et al.
Published: (2025)
by: Hikal, Baraa, et al.
Published: (2025)
LogSigma at SemEval-2026 Task 3: Uncertainty-Weighted Multitask Learning for Dimensional Aspect-Based Sentiment Analysis
by: Hikal, Baraa, et al.
Published: (2026)
by: Hikal, Baraa, et al.
Published: (2026)
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
by: Kochmar, Ekaterina, et al.
Published: (2025)
by: Kochmar, Ekaterina, et al.
Published: (2025)
BD at BEA 2025 Shared Task: MPNet Ensembles for Pedagogical Mistake Identification and Localization in AI Tutor Responses
by: Rohan, Shadman, et al.
Published: (2025)
by: Rohan, Shadman, et al.
Published: (2025)
RETUYT-INCO at BEA 2025 Shared Task: How Far Can Lightweight Models Go in AI-powered Tutor Evaluation?
by: Góngora, Santiago, et al.
Published: (2025)
by: Góngora, Santiago, et al.
Published: (2025)
NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Sakura at BEA 2026 Shared Task 1: What Makes Vocabulary Difficult?
by: Nohejl, Adam, et al.
Published: (2026)
by: Nohejl, Adam, et al.
Published: (2026)
Adaptive Multi-Expert Reasoning via Difficulty-Aware Routing and Uncertainty-Guided Aggregation
by: Ehab, Mohamed, et al.
Published: (2026)
by: Ehab, Mohamed, et al.
Published: (2026)
LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task
by: Leonardelli, Elisa, et al.
Published: (2025)
by: Leonardelli, Elisa, et al.
Published: (2025)
Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
by: Younes, Mohamed T., et al.
Published: (2025)
by: Younes, Mohamed T., et al.
Published: (2025)
LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring
by: Lee, Unggi, et al.
Published: (2026)
by: Lee, Unggi, et al.
Published: (2026)
Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
by: Elgabry, Menna, et al.
Published: (2025)
by: Elgabry, Menna, et al.
Published: (2025)
CoinMath: Harnessing the Power of Coding Instruction for Math LLMs
by: Wei, Chengwei, et al.
Published: (2024)
by: Wei, Chengwei, et al.
Published: (2024)
CAMO: A Class-Aware Minority-Optimized Ensemble for Robust Language Model Evaluation on Imbalanced Data
by: Ehab, Mohamed, et al.
Published: (2026)
by: Ehab, Mohamed, et al.
Published: (2026)
The ADAIO System at the BEA-2023 Shared Task on Generating AI Teacher Responses in Educational Dialogues
by: Adigwe, Adaeze, et al.
Published: (2023)
by: Adigwe, Adaeze, et al.
Published: (2023)
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
RETUYT-INCO at BEA 2026 Shared Task 2: Meta-prompting in Rubric-based Scoring for German
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Large Language Models Approach Expert Pedagogical Quality in Math Tutoring but Differ in Instructional and Linguistic Profiles
by: Abdulsalam, Ramatu Oiza, et al.
Published: (2025)
by: Abdulsalam, Ramatu Oiza, et al.
Published: (2025)
MathBuddy: A Multimodal System for Affective Math Tutoring
by: Kar, Debanjana, et al.
Published: (2025)
by: Kar, Debanjana, et al.
Published: (2025)
Severity-Aware Weighted Loss for Arabic Medical Text Generation
by: Alansary, Ahmed, et al.
Published: (2026)
by: Alansary, Ahmed, et al.
Published: (2026)
MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models
by: Ahmed, Seif, et al.
Published: (2025)
by: Ahmed, Seif, et al.
Published: (2025)
JuniperLiu at CoMeDi Shared Task: Models as Annotators in Lexical Semantics Disagreements
by: Liu, Zhu, et al.
Published: (2024)
by: Liu, Zhu, et al.
Published: (2024)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
by: Yang, Tengchao, et al.
Published: (2025)
by: Yang, Tengchao, et al.
Published: (2025)
Towards Reward Modeling for AI Tutors in Math Mistake Remediation
by: Petukhova, Kseniia, et al.
Published: (2026)
by: Petukhova, Kseniia, et al.
Published: (2026)
A Multi-Layered Large Language Model Framework for Disease Prediction
by: Mohamed, Malak, et al.
Published: (2025)
by: Mohamed, Malak, et al.
Published: (2025)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
Funzac at CoMeDi Shared Task: Modeling Annotator Disagreement from Word-In-Context Perspectives
by: Sarumi, Olufunke O., et al.
Published: (2025)
by: Sarumi, Olufunke O., et al.
Published: (2025)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
by: Tonga, Junior Cedric, et al.
Published: (2025)
by: Tonga, Junior Cedric, et al.
Published: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
by: Tang, Zhengyang, et al.
Published: (2024)
by: Tang, Zhengyang, et al.
Published: (2024)
Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset
by: Zhou, Zhuqian, et al.
Published: (2026)
by: Zhou, Zhuqian, et al.
Published: (2026)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
Similar Items
-
Optimized Quran Passage Retrieval Using an Expanded QA Dataset and Fine-Tuned Language Models
by: Basem, Mohamed, et al.
Published: (2024) -
Few-Shot Prompting for Extractive Quranic QA with Instruction-Tuned LLMs
by: Basem, Mohamed, et al.
Published: (2025) -
MSA at SemEval-2025 Task 3: High Quality Weak Labeling and LLM Ensemble Verification for Multilingual Hallucination Detection
by: Hikal, Baraa, et al.
Published: (2025) -
Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction
by: Basem, Mohamed, et al.
Published: (2025) -
Cross-Language Approach for Quranic QA
by: Oshallah, Islam, et al.
Published: (2025)