Has Automated Essay Scoring Reached Sufficient Accuracy? Deriving Achievable QWK Ceilings from Classical Test Theory
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Uto, Masaki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Has LLM Reached the Scaling Ceiling Yet? Unified Insights into LLM Regularities and Constraints
von: Luo, Charles
Veröffentlicht: (2024)
von: Luo, Charles
Veröffentlicht: (2024)
Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
von: Ormerod, Christopher
Veröffentlicht: (2025)
von: Ormerod, Christopher
Veröffentlicht: (2025)
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
von: Cai, Yida, et al.
Veröffentlicht: (2025)
von: Cai, Yida, et al.
Veröffentlicht: (2025)
Evaluating Austrian A-Level German Essays with Large Language Models for Automated Essay Scoring
von: Kubesch, Jonas, et al.
Veröffentlicht: (2026)
von: Kubesch, Jonas, et al.
Veröffentlicht: (2026)
AI-generated Essays: Characteristics and Implications on Automated Scoring and Academic Integrity
von: Zhong, Yang, et al.
Veröffentlicht: (2024)
von: Zhong, Yang, et al.
Veröffentlicht: (2024)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
Towards Prompt Generalization: Grammar-aware Cross-Prompt Automated Essay Scoring
von: Do, Heejin, et al.
Veröffentlicht: (2025)
von: Do, Heejin, et al.
Veröffentlicht: (2025)
Autoregressive Score Generation for Multi-trait Essay Scoring
von: Do, Heejin, et al.
Veröffentlicht: (2024)
von: Do, Heejin, et al.
Veröffentlicht: (2024)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
von: Chu, SeongYeub, et al.
Veröffentlicht: (2024)
von: Chu, SeongYeub, et al.
Veröffentlicht: (2024)
Leveraging AI Graders for Missing Score Imputation to Achieve Accurate Ability Estimation in Constructed-Response Tests
von: Uto, Masaki, et al.
Veröffentlicht: (2025)
von: Uto, Masaki, et al.
Veröffentlicht: (2025)
LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Models
von: Shibata, Takumi, et al.
Veröffentlicht: (2025)
von: Shibata, Takumi, et al.
Veröffentlicht: (2025)
The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost
von: Frohn, Scott
Veröffentlicht: (2026)
von: Frohn, Scott
Veröffentlicht: (2026)
Teach-to-Reason with Scoring: Self-Explainable Rationale-Driven Multi-Trait Essay Scoring
von: Do, Heejin, et al.
Veröffentlicht: (2025)
von: Do, Heejin, et al.
Veröffentlicht: (2025)
DREsS: Dataset for Rubric-based Essay Scoring on EFL Writing
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
Activations as Features: Probing LLMs for Generalizable Essay Scoring Representations
von: Chi, Jinwei, et al.
Veröffentlicht: (2025)
von: Chi, Jinwei, et al.
Veröffentlicht: (2025)
Improve LLM-based Automatic Essay Scoring with Linguistic Features
von: Hou, Zhaoyi Joey, et al.
Veröffentlicht: (2025)
von: Hou, Zhaoyi Joey, et al.
Veröffentlicht: (2025)
Automatic Essay Scoring and Feedback Generation in Basque Language Learning
von: Azurmendi, Ekhi, et al.
Veröffentlicht: (2025)
von: Azurmendi, Ekhi, et al.
Veröffentlicht: (2025)
Assessing GPTZero's Accuracy in Identifying AI vs. Human-Written Essays
von: Dik, Selin, et al.
Veröffentlicht: (2025)
von: Dik, Selin, et al.
Veröffentlicht: (2025)
Autoregressive Multi-trait Essay Scoring via Reinforcement Learning with Scoring-aware Multiple Rewards
von: Do, Heejin, et al.
Veröffentlicht: (2024)
von: Do, Heejin, et al.
Veröffentlicht: (2024)
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
von: Hallaç, İbrahim Rıza, et al.
Veröffentlicht: (2026)
von: Hallaç, İbrahim Rıza, et al.
Veröffentlicht: (2026)
Automatic Essay Multi-dimensional Scoring with Fine-tuning and Multiple Regression
von: Sun, Kun, et al.
Veröffentlicht: (2024)
von: Sun, Kun, et al.
Veröffentlicht: (2024)
Can Large Language Models Automatically Score Proficiency of Written Essays?
von: Mansour, Watheq, et al.
Veröffentlicht: (2024)
von: Mansour, Watheq, et al.
Veröffentlicht: (2024)
LLM Essay Scoring Under Holistic and Analytic Rubrics: Prompt Effects and Bias
von: Kucia, Filip J., et al.
Veröffentlicht: (2026)
von: Kucia, Filip J., et al.
Veröffentlicht: (2026)
Decision-Level Ordinal Modeling for Multimodal Essay Scoring with Large Language Models
von: Zhang, Han, et al.
Veröffentlicht: (2026)
von: Zhang, Han, et al.
Veröffentlicht: (2026)
Human-AI Collaborative Essay Scoring: A Dual-Process Framework with LLMs
von: Xiao, Changrong, et al.
Veröffentlicht: (2024)
von: Xiao, Changrong, et al.
Veröffentlicht: (2024)
Enhancing Essay Scoring with Adversarial Weights Perturbation and Metric-specific AttentionPooling
von: Huang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Huang, Jiaxin, et al.
Veröffentlicht: (2024)
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion
von: Huang, ShiYing, et al.
Veröffentlicht: (2026)
von: Huang, ShiYing, et al.
Veröffentlicht: (2026)
Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
von: Marín, Javier
Veröffentlicht: (2025)
von: Marín, Javier
Veröffentlicht: (2025)
Diagnosing Spectral Ceilings in Equivariant Neural Force Fields
von: Kim, Hyunmog
Veröffentlicht: (2026)
von: Kim, Hyunmog
Veröffentlicht: (2026)
CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
Transformer-based Joint Modelling for Automatic Essay Scoring and Off-Topic Detection
von: Das, Sourya Dipta, et al.
Veröffentlicht: (2024)
von: Das, Sourya Dipta, et al.
Veröffentlicht: (2024)
Using ChatGPT to Score Essays and Short-Form Constructed Responses
von: Shermis, Mark D.
Veröffentlicht: (2024)
von: Shermis, Mark D.
Veröffentlicht: (2024)
Securing the Floor and Raising the Ceiling: A Merging-based Paradigm for Multi-modal Search Agents
von: Wang, Zhixiang, et al.
Veröffentlicht: (2026)
von: Wang, Zhixiang, et al.
Veröffentlicht: (2026)
Logic-Free Building Automation: Learning the Control of Room Facilities with Wall Switches and Ceiling Camera
von: Ochiai, Hideya, et al.
Veröffentlicht: (2024)
von: Ochiai, Hideya, et al.
Veröffentlicht: (2024)
Enhancing Essay Cohesion Assessment: A Novel Item Response Theory Approach
von: Rosa, Bruno Alexandre, et al.
Veröffentlicht: (2025)
von: Rosa, Bruno Alexandre, et al.
Veröffentlicht: (2025)
Diffusion Denoiser Achievable Analysis for Finite Blocklength Unsourced Random Access
von: Han, Yuming, et al.
Veröffentlicht: (2026)
von: Han, Yuming, et al.
Veröffentlicht: (2026)
An Essay concerning machine understanding
von: Roitblat, Herbert L.
Veröffentlicht: (2024)
von: Roitblat, Herbert L.
Veröffentlicht: (2024)
Assessing the Reliability and Validity of Large Language Models for Automated Assessment of Student Essays in Higher Education
von: Gaggioli, Andrea, et al.
Veröffentlicht: (2025)
von: Gaggioli, Andrea, et al.
Veröffentlicht: (2025)
Ensembling Tabular Foundation Models - A Diversity Ceiling And A Calibration Trap
von: Tanna, Aditya, et al.
Veröffentlicht: (2026)
von: Tanna, Aditya, et al.
Veröffentlicht: (2026)
Diffusion-Inspired Cold Start with Sufficient Prior in Computerized Adaptive Testing
von: Ma, Haiping, et al.
Veröffentlicht: (2024)
von: Ma, Haiping, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Has LLM Reached the Scaling Ceiling Yet? Unified Insights into LLM Regularities and Constraints
von: Luo, Charles
Veröffentlicht: (2024) -
Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
von: Ormerod, Christopher
Veröffentlicht: (2025) -
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
von: Cai, Yida, et al.
Veröffentlicht: (2025) -
Evaluating Austrian A-Level German Essays with Large Language Models for Automated Essay Scoring
von: Kubesch, Jonas, et al.
Veröffentlicht: (2026) -
AI-generated Essays: Characteristics and Implications on Automated Scoring and Academic Integrity
von: Zhong, Yang, et al.
Veröffentlicht: (2024)