FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Seongyeub, Kim, Jongwoo, Yi, Munyong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aspect-Aware MOOC Recommendation in a Heterogeneous Network
by: Chu, Seongyeub, et al.
Published: (2026)
by: Chu, Seongyeub, et al.
Published: (2026)
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
by: Liang, Junhong, et al.
Published: (2026)
by: Liang, Junhong, et al.
Published: (2026)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
by: Chu, SeongYeub, et al.
Published: (2024)
by: Chu, SeongYeub, et al.
Published: (2024)
Exploring LLM Prompting Strategies for Joint Essay Scoring and Feedback Generation
by: Stahl, Maja, et al.
Published: (2024)
by: Stahl, Maja, et al.
Published: (2024)
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
by: Liang, Chumeng, et al.
Published: (2025)
by: Liang, Chumeng, et al.
Published: (2025)
AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents
by: Liang, Yuanzhi, et al.
Published: (2024)
by: Liang, Yuanzhi, et al.
Published: (2024)
RepEval: Effective Text Evaluation with LLM Representation
by: Sheng, Shuqian, et al.
Published: (2024)
by: Sheng, Shuqian, et al.
Published: (2024)
MILE-RefHumEval: A Reference-Free, Multi-Independent LLM Framework for Human-Aligned Evaluation
by: Srun, Nalin, et al.
Published: (2026)
by: Srun, Nalin, et al.
Published: (2026)
EssayCBM: Rubric-Aligned Concept Bottleneck Models for Transparent Essay Grading
by: Chaudhary, Kumar Satvik, et al.
Published: (2025)
by: Chaudhary, Kumar Satvik, et al.
Published: (2025)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
by: Zhou, Yixi, et al.
Published: (2026)
by: Zhou, Yixi, et al.
Published: (2026)
CreativEval: Evaluating Creativity of LLM-Based Hardware Code Generation
by: DeLorenzo, Matthew, et al.
Published: (2024)
by: DeLorenzo, Matthew, et al.
Published: (2024)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
Automatic Essay Scoring and Feedback Generation in Basque Language Learning
by: Azurmendi, Ekhi, et al.
Published: (2025)
by: Azurmendi, Ekhi, et al.
Published: (2025)
HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
by: Liu, Yuxuan, et al.
Published: (2024)
by: Liu, Yuxuan, et al.
Published: (2024)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
by: Zhao, Jiahao, et al.
Published: (2025)
by: Zhao, Jiahao, et al.
Published: (2025)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
by: Shen, Chengyu, et al.
Published: (2026)
by: Shen, Chengyu, et al.
Published: (2026)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
A Better LLM Evaluator for Text Generation: The Impact of Prompt Output Sequencing and Optimization
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
System Report for CCL24-Eval Task 7: Multi-Error Modeling and Fluency-Targeted Pre-training for Chinese Essay Evaluation
by: Zhang, Jingshen, et al.
Published: (2024)
by: Zhang, Jingshen, et al.
Published: (2024)
ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical Judgment
by: Li, Ruochen, et al.
Published: (2025)
by: Li, Ruochen, et al.
Published: (2025)
Pedagogy-R1: Pedagogically-Aligned Reasoning Model with Balanced Educational Benchmark
by: Lee, Unggi, et al.
Published: (2025)
by: Lee, Unggi, et al.
Published: (2025)
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
by: Dejl, Adam, et al.
Published: (2026)
by: Dejl, Adam, et al.
Published: (2026)
RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
by: Wei, Tianjun, et al.
Published: (2025)
by: Wei, Tianjun, et al.
Published: (2025)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
by: Zhou, Lingfeng, et al.
Published: (2025)
by: Zhou, Lingfeng, et al.
Published: (2025)
Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
by: Ko, Jongwoo, et al.
Published: (2025)
by: Ko, Jongwoo, et al.
Published: (2025)
DistiLLM: Towards Streamlined Distillation for Large Language Models
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
Let Me Teach You: Pedagogical Foundations of Feedback for Language Models
by: Borges, Beatriz, et al.
Published: (2023)
by: Borges, Beatriz, et al.
Published: (2023)
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
SpecEval: Evaluating Model Adherence to Behavior Specifications
by: Ahmed, Ahmed, et al.
Published: (2025)
by: Ahmed, Ahmed, et al.
Published: (2025)
PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning
by: Chang, Qikai, et al.
Published: (2026)
by: Chang, Qikai, et al.
Published: (2026)
LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text
by: yunhan, Li, et al.
Published: (2025)
by: yunhan, Li, et al.
Published: (2025)
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
by: Lior, Gili, et al.
Published: (2025)
by: Lior, Gili, et al.
Published: (2025)
ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding
by: Azime, Israel Abebe, et al.
Published: (2024)
by: Azime, Israel Abebe, et al.
Published: (2024)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Can Large Language Models Differentiate Harmful from Argumentative Essays? Steps Toward Ethical Essay Scoring
by: Kim, Hongjin, et al.
Published: (2026)
by: Kim, Hongjin, et al.
Published: (2026)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
by: Song, Seyoung, et al.
Published: (2025)
by: Song, Seyoung, et al.
Published: (2025)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
by: Shi, Taiwei, et al.
Published: (2024)
by: Shi, Taiwei, et al.
Published: (2024)
LLM as a Scorer: The Impact of Output Order on Dialogue Evaluation
by: Chen, Yi-Pei, et al.
Published: (2024)
by: Chen, Yi-Pei, et al.
Published: (2024)
Similar Items
-
Aspect-Aware MOOC Recommendation in a Heterogeneous Network
by: Chu, Seongyeub, et al.
Published: (2026) -
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation
by: Liang, Junhong, et al.
Published: (2026) -
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
by: Chu, SeongYeub, et al.
Published: (2024) -
Exploring LLM Prompting Strategies for Joint Essay Scoring and Feedback Generation
by: Stahl, Maja, et al.
Published: (2024) -
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
by: Liang, Chumeng, et al.
Published: (2025)