Machine-Assisted Grading of Nationwide School-Leaving Essay Exams with LLMs and Statistical NLP
Fuente:
arXiv
Saved in:
| Main Authors: | Karjus, Andres, Allkivi, Kais, Maine, Silvia, Leppik, Katarin, Kruusmaa, Krister, Aruvee, Merilin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards interpretable models for language proficiency assessment: Predicting the CEFR level of Estonian learner texts
by: Allkivi, Kais
Published: (2026)
by: Allkivi, Kais
Published: (2026)
Machine-assisted quantitizing designs: augmenting humanities and social sciences with artificial intelligence
by: Karjus, Andres
Published: (2023)
by: Karjus, Andres
Published: (2023)
LLMs Do Not Grade Essays Like Humans
by: Mathew, Jerin George, et al.
Published: (2026)
by: Mathew, Jerin George, et al.
Published: (2026)
EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
by: Dorkin, Aleksei, et al.
Published: (2026)
by: Dorkin, Aleksei, et al.
Published: (2026)
How well can LLMs Grade Essays in Arabic?
by: Ghazawi, Rayed, et al.
Published: (2025)
by: Ghazawi, Rayed, et al.
Published: (2025)
Autocorrect for Estonian texts: final report from project EKTB25
by: Luhtaru, Agnes, et al.
Published: (2024)
by: Luhtaru, Agnes, et al.
Published: (2024)
Hey AI Can You Grade My Essay?: Automatic Essay Grading
by: Maliha, Maisha, et al.
Published: (2024)
by: Maliha, Maisha, et al.
Published: (2024)
Are Language Models Borrowing-Blind? A Multilingual Evaluation of Loanword Identification across 10 Languages
by: Silva, Mérilin Sousa, et al.
Published: (2025)
by: Silva, Mérilin Sousa, et al.
Published: (2025)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
by: Chu, SeongYeub, et al.
Published: (2024)
by: Chu, SeongYeub, et al.
Published: (2024)
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
by: Zhu, Yunqi, et al.
Published: (2024)
by: Zhu, Yunqi, et al.
Published: (2024)
Advancing NLP Security by Leveraging LLMs as Adversarial Engines
by: Srinivasan, Sudarshan, et al.
Published: (2024)
by: Srinivasan, Sudarshan, et al.
Published: (2024)
Activations as Features: Probing LLMs for Generalizable Essay Scoring Representations
by: Chi, Jinwei, et al.
Published: (2025)
by: Chi, Jinwei, et al.
Published: (2025)
On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMs
by: Calderon, Nitay, et al.
Published: (2024)
by: Calderon, Nitay, et al.
Published: (2024)
Ground Truth Generation for Multilingual Historical NLP using LLMs
by: Gladstone, Clovis, et al.
Published: (2025)
by: Gladstone, Clovis, et al.
Published: (2025)
Designing Reliable LLM-Assisted Rubric Scoring for Constructed Responses: Evidence from Physics Exams
by: Tang, Xiuxiu, et al.
Published: (2026)
by: Tang, Xiuxiu, et al.
Published: (2026)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
by: Liu, Chaoqun, et al.
Published: (2025)
by: Liu, Chaoqun, et al.
Published: (2025)
From Transformers to LLMs: A Systematic Survey of Efficiency Considerations in NLP
by: Ansar, Wazib, et al.
Published: (2024)
by: Ansar, Wazib, et al.
Published: (2024)
Human-AI Collaborative Essay Scoring: A Dual-Process Framework with LLMs
by: Xiao, Changrong, et al.
Published: (2024)
by: Xiao, Changrong, et al.
Published: (2024)
Automated stance detection in complex topics and small languages: the challenging case of immigration in polarizing news media
by: Mets, Mark, et al.
Published: (2023)
by: Mets, Mark, et al.
Published: (2023)
The Ever-Evolving Science Exam
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
Machine Learning-based NLP for Emotion Classification on a Cholera X Dataset
by: Jideani, Paul, et al.
Published: (2024)
by: Jideani, Paul, et al.
Published: (2024)
Decoding AI and Human Authorship: Nuances Revealed Through NLP and Statistical Analysis
by: Akinwande, Mayowa, et al.
Published: (2024)
by: Akinwande, Mayowa, et al.
Published: (2024)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
by: Wang, Minzheng, et al.
Published: (2024)
by: Wang, Minzheng, et al.
Published: (2024)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
by: Mitra, Arindam, et al.
Published: (2024)
by: Mitra, Arindam, et al.
Published: (2024)
Specialists or Generalists? Multi-Agent and Single-Agent LLMs for Essay Grading
by: Idowu, Jamiu Adekunle, et al.
Published: (2026)
by: Idowu, Jamiu Adekunle, et al.
Published: (2026)
On the Evaluation Practices in Multilingual NLP: Can Machine Translation Offer an Alternative to Human Translations?
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
The Pitfalls of Publishing in the Age of LLMs: Strange and Surprising Adventures with a High-Impact NLP Journal
by: Verma, Rakesh M., et al.
Published: (2024)
by: Verma, Rakesh M., et al.
Published: (2024)
Humanity's Last Exam
by: Phan, Long, et al.
Published: (2025)
by: Phan, Long, et al.
Published: (2025)
EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
by: Gao, Fan, et al.
Published: (2025)
by: Gao, Fan, et al.
Published: (2025)
Assisting the Grading of a Handwritten General Chemistry Exam with Artificial Intelligence
by: Cvengros, Jan, et al.
Published: (2025)
by: Cvengros, Jan, et al.
Published: (2025)
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
by: Du, Jiangshu, et al.
Published: (2024)
by: Du, Jiangshu, et al.
Published: (2024)
Evaluating Austrian A-Level German Essays with Large Language Models for Automated Essay Scoring
by: Kubesch, Jonas, et al.
Published: (2026)
by: Kubesch, Jonas, et al.
Published: (2026)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
by: Tahir, Munief Hassan, et al.
Published: (2024)
by: Tahir, Munief Hassan, et al.
Published: (2024)
From Struggle (06-2024) to Mastery (02-2025) LLMs Conquer Advanced Algorithm Exams and Pave the Way for Editorial Generation
by: Dumitran, Adrian Marius, et al.
Published: (2025)
by: Dumitran, Adrian Marius, et al.
Published: (2025)
Evaluating Deduplication Techniques for Economic Research Paper Titles with a Focus on Semantic Similarity using NLP and LLMs
by: You, Doohee, et al.
Published: (2024)
by: You, Doohee, et al.
Published: (2024)
NLP Security and Ethics, in the Wild
by: Lent, Heather, et al.
Published: (2025)
by: Lent, Heather, et al.
Published: (2025)
EssayCBM: Rubric-Aligned Concept Bottleneck Models for Transparent Essay Grading
by: Chaudhary, Kumar Satvik, et al.
Published: (2025)
by: Chaudhary, Kumar Satvik, et al.
Published: (2025)
Socioeconomic factors of national representation in the global film festival circuit: skewed toward the large and wealthy, but small countries can beat the odds
by: Karjus, Andres
Published: (2024)
by: Karjus, Andres
Published: (2024)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
by: Su, Jiamin, et al.
Published: (2025)
by: Su, Jiamin, et al.
Published: (2025)
Reasoning Models Ace the CFA Exams
by: Patel, Jaisal, et al.
Published: (2025)
by: Patel, Jaisal, et al.
Published: (2025)
Similar Items
-
Towards interpretable models for language proficiency assessment: Predicting the CEFR level of Estonian learner texts
by: Allkivi, Kais
Published: (2026) -
Machine-assisted quantitizing designs: augmenting humanities and social sciences with artificial intelligence
by: Karjus, Andres
Published: (2023) -
LLMs Do Not Grade Essays Like Humans
by: Mathew, Jerin George, et al.
Published: (2026) -
EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
by: Dorkin, Aleksei, et al.
Published: (2026) -
How well can LLMs Grade Essays in Arabic?
by: Ghazawi, Rayed, et al.
Published: (2025)