EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Jiamin, Yan, Yibo, Fu, Fangteng, Zhang, Han, Ye, Jingheng, Liu, Xiang, Huo, Jiahao, Zhou, Huiyu, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring
by: Su, Jiamin, et al.
Published: (2025)
by: Su, Jiamin, et al.
Published: (2025)
Decision-Level Ordinal Modeling for Multimodal Essay Scoring with Large Language Models
by: Zhang, Han, et al.
Published: (2026)
by: Zhang, Han, et al.
Published: (2026)
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
by: Yan, Yibo, et al.
Published: (2024)
by: Yan, Yibo, et al.
Published: (2024)
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
by: Yan, Yibo, et al.
Published: (2025)
by: Yan, Yibo, et al.
Published: (2025)
MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model
by: Huo, Jiahao, et al.
Published: (2024)
by: Huo, Jiahao, et al.
Published: (2024)
Evaluating Austrian A-Level German Essays with Large Language Models for Automated Essay Scoring
by: Kubesch, Jonas, et al.
Published: (2026)
by: Kubesch, Jonas, et al.
Published: (2026)
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
by: Cai, Yida, et al.
Published: (2025)
by: Cai, Yida, et al.
Published: (2025)
Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French
by: Wilkens, Rodrigo, et al.
Published: (2026)
by: Wilkens, Rodrigo, et al.
Published: (2026)
MINER: Mining the Underlying Pattern of Modality-Specific Neurons in Multimodal Large Language Models
by: Huang, Kaichen, et al.
Published: (2024)
by: Huang, Kaichen, et al.
Published: (2024)
Long Context Automated Essay Scoring with Language Models
by: Ormerod, Christopher, et al.
Published: (2025)
by: Ormerod, Christopher, et al.
Published: (2025)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
by: Huo, Jiahao, et al.
Published: (2025)
by: Huo, Jiahao, et al.
Published: (2025)
LAILA: A Large Trait-Based Dataset for Arabic Automated Essay Scoring
by: Bashendy, May, et al.
Published: (2025)
by: Bashendy, May, et al.
Published: (2025)
Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
by: Ormerod, Christopher
Published: (2025)
by: Ormerod, Christopher
Published: (2025)
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment
by: Karim, Ahmed, et al.
Published: (2025)
by: Karim, Ahmed, et al.
Published: (2025)
Operationalizing Automated Essay Scoring: A Human-Aware Approach
by: Plasencia-Calaña, Yenisel
Published: (2025)
by: Plasencia-Calaña, Yenisel
Published: (2025)
Exploring the Utilities of the Rationales from Large Language Models to Enhance Automated Essay Scoring
by: Jiao, Hong, et al.
Published: (2025)
by: Jiao, Hong, et al.
Published: (2025)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
by: Chu, SeongYeub, et al.
Published: (2024)
by: Chu, SeongYeub, et al.
Published: (2024)
Can Large Language Models Differentiate Harmful from Argumentative Essays? Steps Toward Ethical Essay Scoring
by: Kim, Hongjin, et al.
Published: (2026)
by: Kim, Hongjin, et al.
Published: (2026)
IELTS Writing Revision Platform with Automated Essay Scoring and Adaptive Feedback
by: Ramancauskas, Titas, et al.
Published: (2025)
by: Ramancauskas, Titas, et al.
Published: (2025)
Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays
by: Hua, Haowei, et al.
Published: (2025)
by: Hua, Haowei, et al.
Published: (2025)
Enhancing Arabic Automated Essay Scoring with Synthetic Data and Error Injection
by: Qwaider, Chatrine, et al.
Published: (2025)
by: Qwaider, Chatrine, et al.
Published: (2025)
AI-generated Essays: Characteristics and Implications on Automated Scoring and Academic Integrity
by: Zhong, Yang, et al.
Published: (2024)
by: Zhong, Yang, et al.
Published: (2024)
AI‐Generated Essays: Characteristics and Implications on Automated Scoring and Academic Integrity
by: Yang Zhong, et al.
Published: (2026)
by: Yang Zhong, et al.
Published: (2026)
Improve LLM-based Automatic Essay Scoring with Linguistic Features
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
by: Yan, Yibo, et al.
Published: (2025)
by: Yan, Yibo, et al.
Published: (2025)
Autoregressive Score Generation for Multi-trait Essay Scoring
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
Assessing the Reliability and Validity of Large Language Models for Automated Assessment of Student Essays in Higher Education
by: Gaggioli, Andrea, et al.
Published: (2025)
by: Gaggioli, Andrea, et al.
Published: (2025)
Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals
by: Wang, Yupei, et al.
Published: (2024)
by: Wang, Yupei, et al.
Published: (2024)
Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
by: Harada, Keno, et al.
Published: (2025)
by: Harada, Keno, et al.
Published: (2025)
Towards Prompt Generalization: Grammar-aware Cross-Prompt Automated Essay Scoring
by: Do, Heejin, et al.
Published: (2025)
by: Do, Heejin, et al.
Published: (2025)
Adversarial Topic-aware Prompt-tuning for Cross-topic Automated Essay Scoring
by: Zhang, Chunyun, et al.
Published: (2025)
by: Zhang, Chunyun, et al.
Published: (2025)
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
by: Yan, Yibo, et al.
Published: (2024)
by: Yan, Yibo, et al.
Published: (2024)
DREsS: Dataset for Rubric-based Essay Scoring on EFL Writing
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Models
by: Shibata, Takumi, et al.
Published: (2025)
by: Shibata, Takumi, et al.
Published: (2025)
Do We Need a Detailed Rubric for Automated Essay Scoring using Large Language Models?
by: Yoshida, Lui
Published: (2025)
by: Yoshida, Lui
Published: (2025)
Automatic Essay Scoring in a Brazilian Scenario
by: Matsuoka, Felipe Akio
Published: (2023)
by: Matsuoka, Felipe Akio
Published: (2023)
Can Large Language Models Automatically Score Proficiency of Written Essays?
by: Mansour, Watheq, et al.
Published: (2024)
by: Mansour, Watheq, et al.
Published: (2024)
Unleashing Large Language Models' Proficiency in Zero-shot Essay Scoring
by: Lee, Sanwoo, et al.
Published: (2024)
by: Lee, Sanwoo, et al.
Published: (2024)
Unveiling the Tapestry of Automated Essay Scoring: A Comprehensive Investigation of Accuracy, Fairness, and Generalizability
by: Yang, Kaixun, et al.
Published: (2024)
by: Yang, Kaixun, et al.
Published: (2024)
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
by: Chakravarty, Abhirup
Published: (2025)
by: Chakravarty, Abhirup
Published: (2025)
Similar Items
-
CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring
by: Su, Jiamin, et al.
Published: (2025) -
Decision-Level Ordinal Modeling for Multimodal Essay Scoring with Large Language Models
by: Zhang, Han, et al.
Published: (2026) -
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
by: Yan, Yibo, et al.
Published: (2024) -
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
by: Yan, Yibo, et al.
Published: (2025) -
MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model
by: Huo, Jiahao, et al.
Published: (2024)