EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Yao, Liang, Rongkeng, Xu, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models (LLMs) as Traffic Control Systems at Urban Intersections: A New Paradigm
by: Masri, Sari, et al.
Published: (2024)
by: Masri, Sari, et al.
Published: (2024)
Agentic Vehicles for Human-Centered Mobility: Definition, Prospects, and Synergistic Co-Development with Vehicle Autonomy
by: Yu, Jiangbo, et al.
Published: (2025)
by: Yu, Jiangbo, et al.
Published: (2025)
Evaluating Large Language Models (LLMs) in Financial NLP: A Comparative Study on Financial Report Analysis
by: Mohsin, Md Talha
Published: (2025)
by: Mohsin, Md Talha
Published: (2025)
Evaluatology: The Science and Engineering of Evaluation
by: Zhan, Jianfeng, et al.
Published: (2024)
by: Zhan, Jianfeng, et al.
Published: (2024)
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
by: Guo, Xingang, et al.
Published: (2025)
by: Guo, Xingang, et al.
Published: (2025)
The Conversational Exam: A Scalable Assessment Design for the AI Era
by: Barba, Lorena A., et al.
Published: (2026)
by: Barba, Lorena A., et al.
Published: (2026)
BiTSA: Leveraging Time Series Foundation Model for Building Energy Analytics
by: Lin, Xiachong, et al.
Published: (2024)
by: Lin, Xiachong, et al.
Published: (2024)
A Systematic Analysis of Declining Medical Safety Messaging in Generative AI Models
by: Sharma, Sonali, et al.
Published: (2025)
by: Sharma, Sonali, et al.
Published: (2025)
TxSum: User-Centered Ethereum Transaction Understanding with Micro-Level Semantic Grounding
by: Peng, Zifan, et al.
Published: (2025)
by: Peng, Zifan, et al.
Published: (2025)
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
by: Padmakumar, Vishakh, et al.
Published: (2026)
by: Padmakumar, Vishakh, et al.
Published: (2026)
AgriTrust: a Federated Semantic Governance Framework for Trusted Agricultural Data Sharing
by: Bergier, Ivan
Published: (2025)
by: Bergier, Ivan
Published: (2025)
Exploring Artificial Intelligence Tutor Teammate Adaptability to Harness Discovery Curiosity and Promote Learning in the Context of Interactive Molecular Dynamics
by: Demir, Mustafa, et al.
Published: (2025)
by: Demir, Mustafa, et al.
Published: (2025)
Gearshift Fellowship: A Next-Generation Neurocomputational Game Platform to Model and Train Human-AI Adaptability
by: Ging-Jehli, Nadja R., et al.
Published: (2025)
by: Ging-Jehli, Nadja R., et al.
Published: (2025)
Digital Inheritance in Web3: A Case Study of Soulbound Tokens and the Social Recovery Pallet within the Polkadot and Kusama Ecosystems
by: Goldston, Justin, et al.
Published: (2023)
by: Goldston, Justin, et al.
Published: (2023)
Intanify AI Platform: Embedded AI for Automated IP Audit and Due Diligence
by: Dorfler, Viktor, et al.
Published: (2025)
by: Dorfler, Viktor, et al.
Published: (2025)
A systematic review and analysis of the viability of virtual reality (VR) in construction work and education
by: Din, Zia Ud, et al.
Published: (2024)
by: Din, Zia Ud, et al.
Published: (2024)
Customized FinGPT Search Agents Using Foundation Models
by: Tian, Felix, et al.
Published: (2024)
by: Tian, Felix, et al.
Published: (2024)
CoQuest: Exploring Research Question Co-Creation with an LLM-based Agent
by: Liu, Yiren, et al.
Published: (2023)
by: Liu, Yiren, et al.
Published: (2023)
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Dialogue
by: Ivey, Jonathan, et al.
Published: (2024)
by: Ivey, Jonathan, et al.
Published: (2024)
Design and Evaluation of Crowd-sourcing Platforms Based on Users Confidence Judgments
by: Ahmadabadi, Samin Nili, et al.
Published: (2022)
by: Ahmadabadi, Samin Nili, et al.
Published: (2022)
Balancing Power and Ethics: A Framework for Addressing Human Rights Concerns in Military AI
by: Islam, Mst Rafia, et al.
Published: (2024)
by: Islam, Mst Rafia, et al.
Published: (2024)
A Methodology for Identifying Evaluation Items for Practical Dialogue Systems Based on Business-Dialogue System Alignment Models
by: Nakano, Mikio, et al.
Published: (2026)
by: Nakano, Mikio, et al.
Published: (2026)
Contact Sensors to Remote Cameras: Quantifying Cardiorespiratory Coupling in High-Altitude Exercise Recovery
by: Tang, Jiankai, et al.
Published: (2025)
by: Tang, Jiankai, et al.
Published: (2025)
Cell-o1: Training LLMs to Solve Single-Cell Reasoning Puzzles with Reinforcement Learning
by: Fang, Yin, et al.
Published: (2025)
by: Fang, Yin, et al.
Published: (2025)
Are Educational Escape Rooms More Effective Than Traditional Lectures for Teaching Software Engineering? A Randomized Controlled Trial
by: Gordillo, Aldo, et al.
Published: (2024)
by: Gordillo, Aldo, et al.
Published: (2024)
NARRA-Gym for Evaluating Interactive Narrative Agents
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Will LLMs be Professional at Fund Investment? DeepFund: A Live Arena Perspective
by: Li, Changlun, et al.
Published: (2025)
by: Li, Changlun, et al.
Published: (2025)
Agora: Teaching the Skill of Consensus-Finding with AI Personas Grounded in Human Voice
by: Ravi, Prerna, et al.
Published: (2026)
by: Ravi, Prerna, et al.
Published: (2026)
Kwame 2.0: Human-in-the-Loop Generative AI Teaching Assistant for Large Scale Online Coding Education in Africa
by: Boateng, George, et al.
Published: (2026)
by: Boateng, George, et al.
Published: (2026)
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
by: Arita, Takaya, et al.
Published: (2025)
by: Arita, Takaya, et al.
Published: (2025)
An Exploration of Effects of Dark Mode on University Students: A Human Computer Interface Analysis
by: Shrestha, Awan, et al.
Published: (2024)
by: Shrestha, Awan, et al.
Published: (2024)
From Sustainable Materials to User-Centered Sustainability: Material Experience in Art Healing
by: Zhang, Yuxin, et al.
Published: (2026)
by: Zhang, Yuxin, et al.
Published: (2026)
Immersive Technologies in Training and Healthcare: From Space Missions to Psychophysiological Research
by: Karpowicz, Barbara, et al.
Published: (2025)
by: Karpowicz, Barbara, et al.
Published: (2025)
How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
by: Bai, Longju, et al.
Published: (2026)
by: Bai, Longju, et al.
Published: (2026)
Synthetic Participatory Planning of Shard Automated Electric Mobility Systems
by: Yu, Jiangbo, et al.
Published: (2024)
by: Yu, Jiangbo, et al.
Published: (2024)
An Offline Mobile Conversational Agent for Mental Health Support: Learning from Emotional Dialogues and Psychological Texts with Student-Centered Evaluation
by: A, Vimaleswar, et al.
Published: (2025)
by: A, Vimaleswar, et al.
Published: (2025)
Real-World Deployment and Evaluation of Kwame for Science, An AI Teaching Assistant for Science Education in West Africa
by: Boateng, George, et al.
Published: (2023)
by: Boateng, George, et al.
Published: (2023)
Auditing and Controlling AI Agent Actions in Spreadsheets
by: Sabouri, Sadra, et al.
Published: (2026)
by: Sabouri, Sadra, et al.
Published: (2026)
LLM Agents for Education: Advances and Applications
by: Chu, Zhendong, et al.
Published: (2025)
by: Chu, Zhendong, et al.
Published: (2025)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
by: Badawi, Abeer, et al.
Published: (2025)
by: Badawi, Abeer, et al.
Published: (2025)
Similar Items
-
Large Language Models (LLMs) as Traffic Control Systems at Urban Intersections: A New Paradigm
by: Masri, Sari, et al.
Published: (2024) -
Agentic Vehicles for Human-Centered Mobility: Definition, Prospects, and Synergistic Co-Development with Vehicle Autonomy
by: Yu, Jiangbo, et al.
Published: (2025) -
Evaluating Large Language Models (LLMs) in Financial NLP: A Comparative Study on Financial Report Analysis
by: Mohsin, Md Talha
Published: (2025) -
Evaluatology: The Science and Engineering of Evaluation
by: Zhan, Jianfeng, et al.
Published: (2024) -
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
by: Guo, Xingang, et al.
Published: (2025)