Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Xuansheng, Saraf, Padmaja Pravin, Lee, Gyeonggeon, Latif, Ehsan, Liu, Ninghao, Zhai, Xiaoming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023)
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023)
Artificial Intelligence Bias on English Language Learners in Automatic Scoring
von: Guo, Shuchen, et al.
Veröffentlicht: (2025)
von: Guo, Shuchen, et al.
Veröffentlicht: (2025)
Efficient Multi-Task Inferencing with a Shared Backbone and Lightweight Task-Specific Adapters for Automatic Scoring
von: Latif, Ehsan, et al.
Veröffentlicht: (2024)
von: Latif, Ehsan, et al.
Veröffentlicht: (2024)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
von: Wang, Yun, et al.
Veröffentlicht: (2026)
von: Wang, Yun, et al.
Veröffentlicht: (2026)
Using GPT-4 to Augment Unbalanced Data for Automatic Scoring
von: Fang, Luyang, et al.
Veröffentlicht: (2023)
von: Fang, Luyang, et al.
Veröffentlicht: (2023)
Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments
von: Latif, Ehsan, et al.
Veröffentlicht: (2023)
von: Latif, Ehsan, et al.
Veröffentlicht: (2023)
Fine-tuning ChatGPT for Automatic Scoring of Written Scientific Explanations in Chinese
von: Yang, Jie, et al.
Veröffentlicht: (2025)
von: Yang, Jie, et al.
Veröffentlicht: (2025)
AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
von: Wang, Yun, et al.
Veröffentlicht: (2026)
von: Wang, Yun, et al.
Veröffentlicht: (2026)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
Human-Centered Design for AI-based Automatically Generated Assessment Reports: A Systematic Review
von: Latif, Ehsan, et al.
Veröffentlicht: (2024)
von: Latif, Ehsan, et al.
Veröffentlicht: (2024)
Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
von: Wu, Xuansheng, et al.
Veröffentlicht: (2024)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2024)
AI Gender Bias, Disparities, and Fairness: Does Training Data Matter?
von: Latif, Ehsan, et al.
Veröffentlicht: (2023)
von: Latif, Ehsan, et al.
Veröffentlicht: (2023)
Advancing Education through Tutoring Systems: A Systematic Literature Review
von: Liu, Vincent, et al.
Veröffentlicht: (2025)
von: Liu, Vincent, et al.
Veröffentlicht: (2025)
Using Generative AI and Multi-Agents to Provide Automatic Feedback
von: Guo, Shuchen, et al.
Veröffentlicht: (2024)
von: Guo, Shuchen, et al.
Veröffentlicht: (2024)
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
von: Mohammadkhani, Ali Ghiasvand
Veröffentlicht: (2024)
von: Mohammadkhani, Ali Ghiasvand
Veröffentlicht: (2024)
Using GPT‐4 to Augment Imbalanced Data for Automatic Scoring
von: Luyang Fang, et al.
Veröffentlicht: (2025)
von: Luyang Fang, et al.
Veröffentlicht: (2025)
Operationalizing Automated Essay Scoring: A Human-Aware Approach
von: Plasencia-Calaña, Yenisel
Veröffentlicht: (2025)
von: Plasencia-Calaña, Yenisel
Veröffentlicht: (2025)
Concept-Guided Chain-of-Thought Prompting for Pairwise Comparison Scoring of Texts with Large Language Models
von: Wu, Patrick Y., et al.
Veröffentlicht: (2023)
von: Wu, Patrick Y., et al.
Veröffentlicht: (2023)
Gemini Pro Defeated by GPT-4V: Evidence from Education
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023)
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
von: Ferrer, Robinson, et al.
Veröffentlicht: (2026)
von: Ferrer, Robinson, et al.
Veröffentlicht: (2026)
Are Large Language Models Good Essay Graders?
von: Kundu, Anindita, et al.
Veröffentlicht: (2024)
von: Kundu, Anindita, et al.
Veröffentlicht: (2024)
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
von: Tam, Zhi Rui, et al.
Veröffentlicht: (2025)
von: Tam, Zhi Rui, et al.
Veröffentlicht: (2025)
The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process
von: Contestabile, Matilde, et al.
Veröffentlicht: (2025)
von: Contestabile, Matilde, et al.
Veröffentlicht: (2025)
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
von: Chakravarty, Abhirup
Veröffentlicht: (2025)
von: Chakravarty, Abhirup
Veröffentlicht: (2025)
Safe in the Future, Dangerous in the Past: Dissecting Temporal and Linguistic Vulnerabilities in LLMs
von: Said, Muhammad Abdullahi, et al.
Veröffentlicht: (2025)
von: Said, Muhammad Abdullahi, et al.
Veröffentlicht: (2025)
AnthroScore: A Computational Linguistic Measure of Anthropomorphism
von: Cheng, Myra, et al.
Veröffentlicht: (2024)
von: Cheng, Myra, et al.
Veröffentlicht: (2024)
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation
von: Neumann, Terrence, et al.
Veröffentlicht: (2024)
von: Neumann, Terrence, et al.
Veröffentlicht: (2024)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
von: Tang, Zeyu, et al.
Veröffentlicht: (2026)
von: Tang, Zeyu, et al.
Veröffentlicht: (2026)
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
von: Vijayaraghavan, Prashanth, et al.
Veröffentlicht: (2025)
von: Vijayaraghavan, Prashanth, et al.
Veröffentlicht: (2025)
Is GPT-4 Alone Sufficient for Automated Essay Scoring?: A Comparative Judgment Approach Based on Rater Cognition
von: Kim, Seungju, et al.
Veröffentlicht: (2024)
von: Kim, Seungju, et al.
Veröffentlicht: (2024)
G-SciEdBERT: A Contextualized LLM for Science Assessment Tasks in German
von: Latif, Ehsan, et al.
Veröffentlicht: (2024)
von: Latif, Ehsan, et al.
Veröffentlicht: (2024)
Cultural Value Differences of LLMs: Prompt, Language, and Model Size
von: Zhong, Qishuai, et al.
Veröffentlicht: (2024)
von: Zhong, Qishuai, et al.
Veröffentlicht: (2024)
Validity Arguments For Constructed Response Scoring Using Generative Artificial Intelligence Applications
von: Casabianca, Jodi M., et al.
Veröffentlicht: (2025)
von: Casabianca, Jodi M., et al.
Veröffentlicht: (2025)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
von: Dong, Wenhan, et al.
Veröffentlicht: (2025)
von: Dong, Wenhan, et al.
Veröffentlicht: (2025)
From Feature-Based Models to Generative AI: Validity Evidence for Constructed Response Scoring
von: Casabianca, Jodi M., et al.
Veröffentlicht: (2026)
von: Casabianca, Jodi M., et al.
Veröffentlicht: (2026)
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2026)
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2026)
Evaluating the Simulation of Human Personality-Driven Susceptibility to Misinformation with LLMs
von: Pratelli, Manuel, et al.
Veröffentlicht: (2025)
von: Pratelli, Manuel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023) -
Artificial Intelligence Bias on English Language Learners in Automatic Scoring
von: Guo, Shuchen, et al.
Veröffentlicht: (2025) -
Efficient Multi-Task Inferencing with a Shared Backbone and Lightweight Task-Specific Adapters for Automatic Scoring
von: Latif, Ehsan, et al.
Veröffentlicht: (2024) -
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
von: Wang, Yun, et al.
Veröffentlicht: (2026) -
Using GPT-4 to Augment Unbalanced Data for Automatic Scoring
von: Fang, Luyang, et al.
Veröffentlicht: (2023)