Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Harada, Keno, Yoshida, Lui, Kojima, Takeshi, Iwasawa, Yusuke, Matsuo, Yutaka |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do We Need a Detailed Rubric for Automated Essay Scoring using Large Language Models?
von: Yoshida, Lui
Veröffentlicht: (2025)
von: Yoshida, Lui
Veröffentlicht: (2025)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025)
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025)
When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following
von: Harada, Keno, et al.
Veröffentlicht: (2025)
von: Harada, Keno, et al.
Veröffentlicht: (2025)
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
von: Kojima, Takeshi, et al.
Veröffentlicht: (2024)
von: Kojima, Takeshi, et al.
Veröffentlicht: (2024)
Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning
von: Takashiro, Shota, et al.
Veröffentlicht: (2024)
von: Takashiro, Shota, et al.
Veröffentlicht: (2024)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
von: Cao, Qi, et al.
Veröffentlicht: (2026)
von: Cao, Qi, et al.
Veröffentlicht: (2026)
The Impact of Example Selection in Few-Shot Prompting on Automated Essay Scoring Using GPT Models
von: Yoshida, Lui
Veröffentlicht: (2024)
von: Yoshida, Lui
Veröffentlicht: (2024)
Dynamic Injection of Entity Knowledge into Dense Retrievers
von: Yamada, Ikuya, et al.
Veröffentlicht: (2025)
von: Yamada, Ikuya, et al.
Veröffentlicht: (2025)
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
von: Yang, Bo, et al.
Veröffentlicht: (2025)
von: Yang, Bo, et al.
Veröffentlicht: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
von: Uchiyama, Fumiya, et al.
Veröffentlicht: (2024)
von: Uchiyama, Fumiya, et al.
Veröffentlicht: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024)
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Long Context Automated Essay Scoring with Language Models
von: Ormerod, Christopher, et al.
Veröffentlicht: (2025)
von: Ormerod, Christopher, et al.
Veröffentlicht: (2025)
TRATES: Trait-Specific Rubric-Assisted Cross-Prompt Essay Scoring
von: Eltanbouly, Sohaila, et al.
Veröffentlicht: (2025)
von: Eltanbouly, Sohaila, et al.
Veröffentlicht: (2025)
Beyond In-Distribution Success: Scaling Curves of CoT Granularity for Language Model Generalization
von: Wang, Ru, et al.
Veröffentlicht: (2025)
von: Wang, Ru, et al.
Veröffentlicht: (2025)
DREsS: Dataset for Rubric-based Essay Scoring on EFL Writing
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
CoReflect: Conversational Evaluation via Co-Evolutionary Simulation and Reflective Rubric Refinement
von: Li, Yunzhe, et al.
Veröffentlicht: (2026)
von: Li, Yunzhe, et al.
Veröffentlicht: (2026)
Understanding Emergent Misalignment via Feature Superposition Geometry
von: Minegishi, Gouki, et al.
Veröffentlicht: (2026)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2026)
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
von: Cai, Yida, et al.
Veröffentlicht: (2025)
von: Cai, Yida, et al.
Veröffentlicht: (2025)
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
IELTS Writing Revision Platform with Automated Essay Scoring and Adaptive Feedback
von: Ramancauskas, Titas, et al.
Veröffentlicht: (2025)
von: Ramancauskas, Titas, et al.
Veröffentlicht: (2025)
Investigating the Multilingual Calibration Effects of Language Model Instruction-Tuning
von: Huang, Jerry, et al.
Veröffentlicht: (2026)
von: Huang, Jerry, et al.
Veröffentlicht: (2026)
EssayCBM: Rubric-Aligned Concept Bottleneck Models for Transparent Essay Grading
von: Chaudhary, Kumar Satvik, et al.
Veröffentlicht: (2025)
von: Chaudhary, Kumar Satvik, et al.
Veröffentlicht: (2025)
LLM Essay Scoring Under Holistic and Analytic Rubrics: Prompt Effects and Bias
von: Kucia, Filip J., et al.
Veröffentlicht: (2026)
von: Kucia, Filip J., et al.
Veröffentlicht: (2026)
Evaluating Austrian A-Level German Essays with Large Language Models for Automated Essay Scoring
von: Kubesch, Jonas, et al.
Veröffentlicht: (2026)
von: Kubesch, Jonas, et al.
Veröffentlicht: (2026)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
von: Wang, Ru, et al.
Veröffentlicht: (2025)
von: Wang, Ru, et al.
Veröffentlicht: (2025)
Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays
von: Hua, Haowei, et al.
Veröffentlicht: (2025)
von: Hua, Haowei, et al.
Veröffentlicht: (2025)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Models
von: Shibata, Takumi, et al.
Veröffentlicht: (2025)
von: Shibata, Takumi, et al.
Veröffentlicht: (2025)
Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French
von: Wilkens, Rodrigo, et al.
Veröffentlicht: (2026)
von: Wilkens, Rodrigo, et al.
Veröffentlicht: (2026)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
von: Wang, Yun, et al.
Veröffentlicht: (2026)
von: Wang, Yun, et al.
Veröffentlicht: (2026)
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
von: Watanabe, Yusuke, et al.
Veröffentlicht: (2026)
von: Watanabe, Yusuke, et al.
Veröffentlicht: (2026)
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
von: Liu, Junyu, et al.
Veröffentlicht: (2026)
von: Liu, Junyu, et al.
Veröffentlicht: (2026)
Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
von: Ormerod, Christopher
Veröffentlicht: (2025)
von: Ormerod, Christopher
Veröffentlicht: (2025)
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
von: Chakravarty, Abhirup
Veröffentlicht: (2025)
von: Chakravarty, Abhirup
Veröffentlicht: (2025)
Enhancing Arabic Automated Essay Scoring with Synthetic Data and Error Injection
von: Qwaider, Chatrine, et al.
Veröffentlicht: (2025)
von: Qwaider, Chatrine, et al.
Veröffentlicht: (2025)
Operationalizing Automated Essay Scoring: A Human-Aware Approach
von: Plasencia-Calaña, Yenisel
Veröffentlicht: (2025)
von: Plasencia-Calaña, Yenisel
Veröffentlicht: (2025)
Ähnliche Einträge
-
Do We Need a Detailed Rubric for Automated Essay Scoring using Large Language Models?
von: Yoshida, Lui
Veröffentlicht: (2025) -
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025) -
When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following
von: Harada, Keno, et al.
Veröffentlicht: (2025) -
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
von: Kojima, Takeshi, et al.
Veröffentlicht: (2024) -
Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning
von: Takashiro, Shota, et al.
Veröffentlicht: (2024)