RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
Fuente:
arXiv
Saved in:
| Main Author: | Mohammadkhani, Ali Ghiasvand |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gap-Filling Prompting Enhances Code-Assisted Mathematical Reasoning
by: Mohammadkhani, Mohammad Ghiasvand
Published: (2024)
by: Mohammadkhani, Mohammad Ghiasvand
Published: (2024)
Operationalizing Automated Essay Scoring: A Human-Aware Approach
by: Plasencia-Calaña, Yenisel
Published: (2025)
by: Plasencia-Calaña, Yenisel
Published: (2025)
Zero-Shot Learning and Key Points Are All You Need for Automated Fact-Checking
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2024)
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2024)
An investigation of structures responsible for gender bias in BERT and DistilBERT
by: Leteno, Thibaud, et al.
Published: (2024)
by: Leteno, Thibaud, et al.
Published: (2024)
EmoScan: Automatic Screening of Depression Symptoms in Romanized Sinhala Tweets
by: Hewapathirana, Jayathi, et al.
Published: (2024)
by: Hewapathirana, Jayathi, et al.
Published: (2024)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
by: An, Bang, et al.
Published: (2024)
by: An, Bang, et al.
Published: (2024)
Checklist Engineering Empowers Multilingual LLM Judges
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2025)
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2025)
Transformer-based Joint Modelling for Automatic Essay Scoring and Off-Topic Detection
by: Das, Sourya Dipta, et al.
Published: (2024)
by: Das, Sourya Dipta, et al.
Published: (2024)
Augmenting Human-Annotated Training Data with Large Language Model Generation and Distillation in Open-Response Assessment
by: Borchers, Conrad, et al.
Published: (2025)
by: Borchers, Conrad, et al.
Published: (2025)
Reasoning Models Generate Societies of Thought
by: Kim, Junsol, et al.
Published: (2026)
by: Kim, Junsol, et al.
Published: (2026)
GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
by: Dumitran, Adrian-Marius, et al.
Published: (2025)
by: Dumitran, Adrian-Marius, et al.
Published: (2025)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
Social Determinants of Health Prediction for ICD-9 Code with Reasoning Models
by: Khan, Sharim, et al.
Published: (2025)
by: Khan, Sharim, et al.
Published: (2025)
Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation
by: Franceschelli, Giorgio, et al.
Published: (2025)
by: Franceschelli, Giorgio, et al.
Published: (2025)
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
by: Patil, Avinash, et al.
Published: (2025)
by: Patil, Avinash, et al.
Published: (2025)
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
by: Hallaç, İbrahim Rıza, et al.
Published: (2026)
by: Hallaç, İbrahim Rıza, et al.
Published: (2026)
Phrase-Level Adversarial Training for Mitigating Bias in Neural Network-based Automatic Essay Scoring
by: Philip, Haddad, et al.
Published: (2024)
by: Philip, Haddad, et al.
Published: (2024)
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
by: Chakravarty, Abhirup
Published: (2025)
by: Chakravarty, Abhirup
Published: (2025)
Towards Explainable Evaluation Metrics for Machine Translation
by: Leiter, Christoph, et al.
Published: (2023)
by: Leiter, Christoph, et al.
Published: (2023)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
by: Anzenberg, Eitan, et al.
Published: (2025)
by: Anzenberg, Eitan, et al.
Published: (2025)
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment
by: Karim, Ahmed, et al.
Published: (2025)
by: Karim, Ahmed, et al.
Published: (2025)
TransGAT: Transformer-Based Graph Neural Networks for Multi-Dimensional Automated Essay Scoring
by: Aljuaid, Hind, et al.
Published: (2025)
by: Aljuaid, Hind, et al.
Published: (2025)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
by: Hamman, Faisal, et al.
Published: (2025)
by: Hamman, Faisal, et al.
Published: (2025)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?
by: de Wynter, Adrian, et al.
Published: (2024)
by: de Wynter, Adrian, et al.
Published: (2024)
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
by: Caraeni, Adriana, et al.
Published: (2024)
by: Caraeni, Adriana, et al.
Published: (2024)
PRSM: A Measure to Evaluate CLIP's Robustness Against Paraphrases
by: Schlegel, Udo, et al.
Published: (2025)
by: Schlegel, Udo, et al.
Published: (2025)
Generalization in Healthcare AI: Evaluation of a Clinical Large Language Model
by: Rahman, Salman, et al.
Published: (2024)
by: Rahman, Salman, et al.
Published: (2024)
Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench
by: Friedrich, Felix, et al.
Published: (2025)
by: Friedrich, Felix, et al.
Published: (2025)
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
by: Wei, Shou'ang, et al.
Published: (2025)
by: Wei, Shou'ang, et al.
Published: (2025)
FairPair: A Robust Evaluation of Biases in Language Models through Paired Perturbations
by: Dwivedi-Yu, Jane, et al.
Published: (2024)
by: Dwivedi-Yu, Jane, et al.
Published: (2024)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays
by: Hua, Haowei, et al.
Published: (2025)
by: Hua, Haowei, et al.
Published: (2025)
Evaluating the Effectiveness of XAI Techniques for Encoder-Based Language Models
by: Mersha, Melkamu Abay, et al.
Published: (2025)
by: Mersha, Melkamu Abay, et al.
Published: (2025)
Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches
by: Azime, Israel Abebe, et al.
Published: (2025)
by: Azime, Israel Abebe, et al.
Published: (2025)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
by: Kohankhaki, Farnaz, et al.
Published: (2024)
by: Kohankhaki, Farnaz, et al.
Published: (2024)
ChatGPT Needs SPADE (Sustainability, PrivAcy, Digital divide, and Ethics) Evaluation: A Review
by: Khowaja, Sunder Ali, et al.
Published: (2023)
by: Khowaja, Sunder Ali, et al.
Published: (2023)
Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models
by: Feng, Duanyu, et al.
Published: (2023)
by: Feng, Duanyu, et al.
Published: (2023)
DAIC-WOZ: On the Validity of Using the Therapist's prompts in Automatic Depression Detection from Clinical Interviews
by: Burdisso, Sergio, et al.
Published: (2024)
by: Burdisso, Sergio, et al.
Published: (2024)
Similar Items
-
Gap-Filling Prompting Enhances Code-Assisted Mathematical Reasoning
by: Mohammadkhani, Mohammad Ghiasvand
Published: (2024) -
Operationalizing Automated Essay Scoring: A Human-Aware Approach
by: Plasencia-Calaña, Yenisel
Published: (2025) -
Zero-Shot Learning and Key Points Are All You Need for Automated Fact-Checking
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2024) -
An investigation of structures responsible for gender bias in BERT and DistilBERT
by: Leteno, Thibaud, et al.
Published: (2024) -
EmoScan: Automatic Screening of Depression Symptoms in Romanized Sinhala Tweets
by: Hewapathirana, Jayathi, et al.
Published: (2024)