ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Shou'ang, Wang, Xinyun, Bi, Shuzhen, Chen, Jian, Li, Ruijia, Jiang, Bo, Lin, Xin, Zhang, Min, Song, Yu, Li, BingDong, Zhou, Aimin, Hao, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content
by: Bi, Shuzhen, et al.
Published: (2026)
by: Bi, Shuzhen, et al.
Published: (2026)
AutoSynth: Automated Workflow Optimization for High-Quality Synthetic Dataset Generation via Monte Carlo Tree Search
by: Bi, Shuzhen, et al.
Published: (2025)
by: Bi, Shuzhen, et al.
Published: (2025)
OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
by: Zhang, Min, et al.
Published: (2025)
by: Zhang, Min, et al.
Published: (2025)
DiaCDM: Cognitive Diagnosis in Teacher-Student Dialogues using the Initiation-Response-Evaluation Framework
by: Jia, Rui, et al.
Published: (2025)
by: Jia, Rui, et al.
Published: (2025)
CoupNeRF: Property‐aware Neural Radiance Fields for Multi‐Material Coupled Scenario Reconstruction
by: Jin Li, et al.
Published: (2024)
by: Jin Li, et al.
Published: (2024)
EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus
by: Wei, Shouang, et al.
Published: (2025)
by: Wei, Shouang, et al.
Published: (2025)
Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction
by: Bi, Shuzhen, et al.
Published: (2026)
by: Bi, Shuzhen, et al.
Published: (2026)
EduResearchBench: A Hierarchical Atomic Task Decomposition Benchmark for Full-Lifecycle Educational Research
by: Yue, Houping, et al.
Published: (2026)
by: Yue, Houping, et al.
Published: (2026)
Agentic Workflow for Education: Concepts and Applications
by: Jiang, Yuan-Hao, et al.
Published: (2025)
by: Jiang, Yuan-Hao, et al.
Published: (2025)
AI Agent for Education: von Neumann Multi-Agent System Framework
by: Jiang, Yuan-Hao, et al.
Published: (2024)
by: Jiang, Yuan-Hao, et al.
Published: (2024)
Thinking in Graphs with CoMAP: A Shared Visual Workspace for Designing Project-Based Learning
by: Li, Ruijia, et al.
Published: (2026)
by: Li, Ruijia, et al.
Published: (2026)
Scaling Laws for Educational AI Agents
by: Wu, Mengsong, et al.
Published: (2026)
by: Wu, Mengsong, et al.
Published: (2026)
How Real Is AI Tutoring? Comparing Simulated and Human Dialogues in One-on-One Instruction
by: Li, Ruijia, et al.
Published: (2025)
by: Li, Ruijia, et al.
Published: (2025)
A First Look at Kolmogorov-Arnold Networks in Surrogate-assisted Evolutionary Algorithms
by: Hao, Hao, et al.
Published: (2024)
by: Hao, Hao, et al.
Published: (2024)
Relation Reasoning with LLMs in Expensive Optimization
by: Lu, Ye, et al.
Published: (2026)
by: Lu, Ye, et al.
Published: (2026)
Un-evaluated Solutions May Be Valuable in Expensive Optimization
by: Hao, Hao, et al.
Published: (2024)
by: Hao, Hao, et al.
Published: (2024)
Large Language Models as Surrogate Models in Evolutionary Algorithms: A Preliminary Study
by: Hao, Hao, et al.
Published: (2024)
by: Hao, Hao, et al.
Published: (2024)
Model Uncertainty in Evolutionary Optimization and Bayesian Optimization: A Comparative Analysis
by: Hao, Hao, et al.
Published: (2024)
by: Hao, Hao, et al.
Published: (2024)
PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model
by: Gao, Mingju, et al.
Published: (2025)
by: Gao, Mingju, et al.
Published: (2025)
Edu-MMBias: A Three-Tier Multimodal Benchmark for Auditing Social Bias in Vision-Language Models under Educational Contexts
by: Li, Ruijia, et al.
Published: (2026)
by: Li, Ruijia, et al.
Published: (2026)
ORCDF: An Oversmoothing-Resistant Cognitive Diagnosis Framework for Student Learning in Online Education Systems
by: Qian, Hong, et al.
Published: (2024)
by: Qian, Hong, et al.
Published: (2024)
DOP: Diagnostic-Oriented Prompting for Large Language Models in Mathematical Correction
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
CRUISE: Cooperative Reconstruction and Editing in V2X Scenarios using Gaussian Splatting
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis
by: Gao, Huan-ang, et al.
Published: (2024)
by: Gao, Huan-ang, et al.
Published: (2024)
xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation
by: Yu, Qingchen, et al.
Published: (2024)
by: Yu, Qingchen, et al.
Published: (2024)
SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation
by: Jiang, Lai, et al.
Published: (2025)
by: Jiang, Lai, et al.
Published: (2025)
HTDC: Hesitation-Triggered Differential Calibration for Mitigating Hallucination in Large Vision-Language Models
by: Liu, Xinyun
Published: (2026)
by: Liu, Xinyun
Published: (2026)
EduVerse: A User-Defined Multi-Agent Simulation Space for Education Scenario
by: Ma, Yiping, et al.
Published: (2025)
by: Ma, Yiping, et al.
Published: (2025)
Colorectal Polyp Segmentation in the Deep Learning Era: A Comprehensive Survey
by: Wu, Zhenyu, et al.
Published: (2024)
by: Wu, Zhenyu, et al.
Published: (2024)
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
Agentic AI for Particle-Based Simulation: Automating SPH Workflows for Debris Flow Modeling
by: Zhang, Danrong, et al.
Published: (2026)
by: Zhang, Danrong, et al.
Published: (2026)
Dual-frame Fluid Motion Estimation with Test-time Optimization and Zero-divergence Loss
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
by: Mazeika, Mantas, et al.
Published: (2024)
by: Mazeika, Mantas, et al.
Published: (2024)
It's Morphing Time: Unleashing the Potential of Multiple LLMs via Multi-objective Optimization
by: Li, Bingdong, et al.
Published: (2024)
by: Li, Bingdong, et al.
Published: (2024)
A Generic and Automated Methodology to Simulate Melting Point
by: Dai, Fu-Zhi, et al.
Published: (2024)
by: Dai, Fu-Zhi, et al.
Published: (2024)
APE-Bench: Evaluating Automated Proof Engineering for Formal Math Libraries
by: Xin, Huajian, et al.
Published: (2025)
by: Xin, Huajian, et al.
Published: (2025)
RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework
by: Zhu, Kunlun, et al.
Published: (2024)
by: Zhu, Kunlun, et al.
Published: (2024)
SmartRefine: A Scenario-Adaptive Refinement Framework for Efficient Motion Prediction
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
ChatSUMO: Large Language Model for Automating Traffic Scenario Generation in Simulation of Urban MObility
by: Li, Shuyang, et al.
Published: (2024)
by: Li, Shuyang, et al.
Published: (2024)
When LLMs Learn to be Students: The SOEI Framework for Modeling and Evaluating Virtual Student Agents in Educational Interaction
by: Ma, Yiping, et al.
Published: (2024)
by: Ma, Yiping, et al.
Published: (2024)
Similar Items
-
EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content
by: Bi, Shuzhen, et al.
Published: (2026) -
AutoSynth: Automated Workflow Optimization for High-Quality Synthetic Dataset Generation via Monte Carlo Tree Search
by: Bi, Shuzhen, et al.
Published: (2025) -
OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
by: Zhang, Min, et al.
Published: (2025) -
DiaCDM: Cognitive Diagnosis in Teacher-Student Dialogues using the Initiation-Response-Evaluation Framework
by: Jia, Rui, et al.
Published: (2025) -
CoupNeRF: Property‐aware Neural Radiance Fields for Multi‐Material Coupled Scenario Reconstruction
by: Jin Li, et al.
Published: (2024)