GRADE: Generalizable Reasoning-Aware Dialogue Evaluation for AI Tutors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bhalerao, Parth, Chang, Jeromy, Chou, David, Ignat, Oana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
von: Bhalerao, Parth, et al.
Veröffentlicht: (2026)
von: Bhalerao, Parth, et al.
Veröffentlicht: (2026)
MAVEN A Multi-Agent Framework for Multicultural Text-to-Video Generation
von: Li, Shuowei, et al.
Veröffentlicht: (2026)
von: Li, Shuowei, et al.
Veröffentlicht: (2026)
When Cultures Meet: Multicultural Text-to-Image Generation
von: Bhalerao, Parth, et al.
Veröffentlicht: (2025)
von: Bhalerao, Parth, et al.
Veröffentlicht: (2025)
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
von: Zhao, Yuming, et al.
Veröffentlicht: (2026)
von: Zhao, Yuming, et al.
Veröffentlicht: (2026)
Cross-cultural Inspiration Detection and Analysis in Real and LLM-generated Social Media Data
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
Interpretable Difficulty-Aware Knowledge Tracing in Tutor-Student Dialogues
von: Huang, Shuyan, et al.
Veröffentlicht: (2026)
von: Huang, Shuyan, et al.
Veröffentlicht: (2026)
Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues
von: Duan, Zhangqi, et al.
Veröffentlicht: (2026)
von: Duan, Zhangqi, et al.
Veröffentlicht: (2026)
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data
von: Mori, Shinka, et al.
Veröffentlicht: (2024)
von: Mori, Shinka, et al.
Veröffentlicht: (2024)
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
von: Maurya, Kaushal Kumar, et al.
Veröffentlicht: (2024)
von: Maurya, Kaushal Kumar, et al.
Veröffentlicht: (2024)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
von: Nwatu, Joan, et al.
Veröffentlicht: (2024)
von: Nwatu, Joan, et al.
Veröffentlicht: (2024)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
von: Bai, Longju, et al.
Veröffentlicht: (2024)
von: Bai, Longju, et al.
Veröffentlicht: (2024)
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
Simulated Students in Tutoring Dialogues: Substance or Illusion?
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2026)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2026)
Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
von: Ignat, Oana, et al.
Veröffentlicht: (2023)
von: Ignat, Oana, et al.
Veröffentlicht: (2023)
Toward Automated Qualitative Analysis: Leveraging Large Language Models for Tutoring Dialogue Evaluation
von: Gu, Megan, et al.
Veröffentlicht: (2025)
von: Gu, Megan, et al.
Veröffentlicht: (2025)
Exploring LLMs for Predicting Tutor Strategy and Student Outcomes in Dialogues
von: Ikram, Fareya, et al.
Veröffentlicht: (2025)
von: Ikram, Fareya, et al.
Veröffentlicht: (2025)
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
von: Nwatu, Joan, et al.
Veröffentlicht: (2025)
von: Nwatu, Joan, et al.
Veröffentlicht: (2025)
Ensembling Large Language Models to Characterize Affective Dynamics in Student-AI Tutor Dialogues
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
von: Hazra, Rima, et al.
Veröffentlicht: (2026)
von: Hazra, Rima, et al.
Veröffentlicht: (2026)
Training LLM-based Tutors to Improve Student Learning Outcomes in Dialogues
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2025)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2025)
Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors
von: Wang, Jian, et al.
Veröffentlicht: (2025)
von: Wang, Jian, et al.
Veröffentlicht: (2025)
Exploring Knowledge Tracing in Tutor-Student Dialogues using LLMs
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2024)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2024)
How Real Is AI Tutoring? Comparing Simulated and Human Dialogues in One-on-One Instruction
von: Li, Ruijia, et al.
Veröffentlicht: (2025)
von: Li, Ruijia, et al.
Veröffentlicht: (2025)
PIIvot: A Lightweight NLP Anonymization Framework for Question-Anchored Tutoring Dialogues
von: Zent, Matthew, et al.
Veröffentlicht: (2025)
von: Zent, Matthew, et al.
Veröffentlicht: (2025)
Pedagogy-driven Evaluation of Generative AI-powered Intelligent Tutoring Systems
von: Maurya, Kaushal Kumar, et al.
Veröffentlicht: (2025)
von: Maurya, Kaushal Kumar, et al.
Veröffentlicht: (2025)
Automated Classification of Tutors' Dialogue Acts Using Generative AI: A Case Study Using the CIMA Corpus
von: He, Liqun, et al.
Veröffentlicht: (2025)
von: He, Liqun, et al.
Veröffentlicht: (2025)
Misconception Diagnosis From Student-Tutor Dialogue: Generate, Retrieve, Rerank
von: Mitton, Joshua, et al.
Veröffentlicht: (2026)
von: Mitton, Joshua, et al.
Veröffentlicht: (2026)
Leveraging LLMs to Assess Tutor Moves in Real-Life Dialogues: A Feasibility Study
von: Thomas, Danielle R., et al.
Veröffentlicht: (2025)
von: Thomas, Danielle R., et al.
Veröffentlicht: (2025)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
von: Jin, Keyan, et al.
Veröffentlicht: (2025)
von: Jin, Keyan, et al.
Veröffentlicht: (2025)
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
von: Badola, Kartikeya, et al.
Veröffentlicht: (2025)
von: Badola, Kartikeya, et al.
Veröffentlicht: (2025)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
AITutor-EvalKit: Exploring the Capabilities of AI Tutors
von: Naeem, Numaan, et al.
Veröffentlicht: (2025)
von: Naeem, Numaan, et al.
Veröffentlicht: (2025)
Exploring Personality-Aware Interactions in Salesperson Dialogue Agents
von: Cheng, Sijia, et al.
Veröffentlicht: (2025)
von: Cheng, Sijia, et al.
Veröffentlicht: (2025)
Letting Tutor Personas "Speak Up" for LLMs: Learning Steering Vectors from Dialogue via Preference Optimization
von: Lee, Jaewook, et al.
Veröffentlicht: (2026)
von: Lee, Jaewook, et al.
Veröffentlicht: (2026)
ProtoReasoning: Prototypes as the Foundation for Generalizable Reasoning in LLMs
von: He, Feng, et al.
Veröffentlicht: (2025)
von: He, Feng, et al.
Veröffentlicht: (2025)
Towards Reward Modeling for AI Tutors in Math Mistake Remediation
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2026)
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2026)
MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors
von: Hikal, Baraa, et al.
Veröffentlicht: (2025)
von: Hikal, Baraa, et al.
Veröffentlicht: (2025)
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
von: Yang, Tengchao, et al.
Veröffentlicht: (2025)
von: Yang, Tengchao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
von: Bhalerao, Parth, et al.
Veröffentlicht: (2026) -
MAVEN A Multi-Agent Framework for Multicultural Text-to-Video Generation
von: Li, Shuowei, et al.
Veröffentlicht: (2026) -
When Cultures Meet: Multicultural Text-to-Image Generation
von: Bhalerao, Parth, et al.
Veröffentlicht: (2025) -
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
von: Ignat, Oana, et al.
Veröffentlicht: (2024) -
Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
von: Zhao, Yuming, et al.
Veröffentlicht: (2026)