Is GPT-4 Alone Sufficient for Automated Essay Scoring?: A Comparative Judgment Approach Based on Rater Cognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Seungju, Jo, Meounggun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Operationalizing Automated Essay Scoring: A Human-Aware Approach
di: Plasencia-Calaña, Yenisel
Pubblicazione: (2025)
di: Plasencia-Calaña, Yenisel
Pubblicazione: (2025)
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
di: Chakravarty, Abhirup
Pubblicazione: (2025)
di: Chakravarty, Abhirup
Pubblicazione: (2025)
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
di: Mohammadkhani, Ali Ghiasvand
Pubblicazione: (2024)
di: Mohammadkhani, Ali Ghiasvand
Pubblicazione: (2024)
Using GPT-4 to Augment Unbalanced Data for Automatic Scoring
di: Fang, Luyang, et al.
Pubblicazione: (2023)
di: Fang, Luyang, et al.
Pubblicazione: (2023)
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals
di: Ehara, Yo
Pubblicazione: (2026)
di: Ehara, Yo
Pubblicazione: (2026)
Better Call GPT, Comparing Large Language Models Against Lawyers
di: Martin, Lauren, et al.
Pubblicazione: (2024)
di: Martin, Lauren, et al.
Pubblicazione: (2024)
DeID-GPT: Zero-shot Medical Text De-Identification by GPT-4
di: Liu, Zhengliang, et al.
Pubblicazione: (2023)
di: Liu, Zhengliang, et al.
Pubblicazione: (2023)
The Impact of Example Selection in Few-Shot Prompting on Automated Essay Scoring Using GPT Models
di: Yoshida, Lui
Pubblicazione: (2024)
di: Yoshida, Lui
Pubblicazione: (2024)
How ChatGPT Changed the Media's Narratives on AI: A Semi-Automated Narrative Analysis Through Frame Semantics
di: Ryazanov, Igor, et al.
Pubblicazione: (2024)
di: Ryazanov, Igor, et al.
Pubblicazione: (2024)
Disentangling Learning from Judgment: Representation Learning for Open Response Analytics
di: Borchers, Conrad, et al.
Pubblicazione: (2025)
di: Borchers, Conrad, et al.
Pubblicazione: (2025)
Automated Item Neutralization for Non-Cognitive Scales: A Large Language Model Approach to Reducing Social-Desirability Bias
di: Wu, Sirui, et al.
Pubblicazione: (2025)
di: Wu, Sirui, et al.
Pubblicazione: (2025)
Unfair TOS: An Automated Approach using Customized BERT
di: Akash, Bathini Sai, et al.
Pubblicazione: (2024)
di: Akash, Bathini Sai, et al.
Pubblicazione: (2024)
Prompting ChatGPT for Translation: A Comparative Analysis of Translation Brief and Persona Prompts
di: He, Sui
Pubblicazione: (2024)
di: He, Sui
Pubblicazione: (2024)
Math anxiety and associative knowledge structure are entwined in psychology students but not in Large Language Models like GPT-3.5 and GPT-4o
di: Ciringione, Luciana, et al.
Pubblicazione: (2025)
di: Ciringione, Luciana, et al.
Pubblicazione: (2025)
LAILA: A Large Trait-Based Dataset for Arabic Automated Essay Scoring
di: Bashendy, May, et al.
Pubblicazione: (2025)
di: Bashendy, May, et al.
Pubblicazione: (2025)
GLARE: Agentic Reasoning for Legal Judgment Prediction
di: Yang, Xinyu, et al.
Pubblicazione: (2025)
di: Yang, Xinyu, et al.
Pubblicazione: (2025)
AppellateGen: A Benchmark for Appellate Legal Judgment Generation
di: Yang, Hongkun, et al.
Pubblicazione: (2026)
di: Yang, Hongkun, et al.
Pubblicazione: (2026)
From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
di: Li, Yang, et al.
Pubblicazione: (2025)
di: Li, Yang, et al.
Pubblicazione: (2025)
ChatGPT for President! Presupposed content in politicians versus GPT-generated texts
di: Garassino, Davide, et al.
Pubblicazione: (2025)
di: Garassino, Davide, et al.
Pubblicazione: (2025)
Are Large Language Models Good Essay Graders?
di: Kundu, Anindita, et al.
Pubblicazione: (2024)
di: Kundu, Anindita, et al.
Pubblicazione: (2024)
Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading
di: Ding, Ming, et al.
Pubblicazione: (2025)
di: Ding, Ming, et al.
Pubblicazione: (2025)
Long Context Automated Essay Scoring with Language Models
di: Ormerod, Christopher, et al.
Pubblicazione: (2025)
di: Ormerod, Christopher, et al.
Pubblicazione: (2025)
A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2026)
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2026)
Towards Automated Situation Awareness: A RAG-Based Framework for Peacebuilding Reports
di: Nemkova, Poli A., et al.
Pubblicazione: (2025)
di: Nemkova, Poli A., et al.
Pubblicazione: (2025)
Predicting challenge moments from students' discourse: A comparison of GPT-4 to two traditional natural language processing approaches
di: Suraworachet, Wannapon, et al.
Pubblicazione: (2024)
di: Suraworachet, Wannapon, et al.
Pubblicazione: (2024)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
di: Wu, Xuansheng, et al.
Pubblicazione: (2024)
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
di: Wang, Bing, et al.
Pubblicazione: (2026)
di: Wang, Bing, et al.
Pubblicazione: (2026)
Legal Fact Prediction: The Missing Piece in Legal Judgment Prediction
di: Liu, Junkai, et al.
Pubblicazione: (2024)
di: Liu, Junkai, et al.
Pubblicazione: (2024)
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
di: Fernandes, Gustavo Lúcius, et al.
Pubblicazione: (2026)
di: Fernandes, Gustavo Lúcius, et al.
Pubblicazione: (2026)
Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs
di: Sun, Huaman, et al.
Pubblicazione: (2023)
di: Sun, Huaman, et al.
Pubblicazione: (2023)
RogueGPT: dis-ethical tuning transforms ChatGPT4 into a Rogue AI in 158 Words
di: Buscemi, Alessio, et al.
Pubblicazione: (2024)
di: Buscemi, Alessio, et al.
Pubblicazione: (2024)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
Big5PersonalityEssays: Introducing a Novel Synthetic Generated Dataset Consisting of Short State-of-Consciousness Essays Annotated Based on the Five Factor Model of Personality
di: Floroiu, Iustin
Pubblicazione: (2024)
di: Floroiu, Iustin
Pubblicazione: (2024)
GPT-4V Cannot Generate Radiology Reports Yet
di: Jiang, Yuyang, et al.
Pubblicazione: (2024)
di: Jiang, Yuyang, et al.
Pubblicazione: (2024)
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
di: Caraeni, Adriana, et al.
Pubblicazione: (2024)
di: Caraeni, Adriana, et al.
Pubblicazione: (2024)
Does GPT-4 surpass human performance in linguistic pragmatics?
di: Bojic, Ljubisa, et al.
Pubblicazione: (2023)
di: Bojic, Ljubisa, et al.
Pubblicazione: (2023)
Analyzing the Performance of ChatGPT in Cardiology and Vascular Pathologies
di: Hariri, Walid
Pubblicazione: (2023)
di: Hariri, Walid
Pubblicazione: (2023)
Artificial Intelligence Driven Course Generation: A Case Study Using ChatGPT
di: Rouabhia, Djaber
Pubblicazione: (2024)
di: Rouabhia, Djaber
Pubblicazione: (2024)
Politicians vs ChatGPT. A study of presuppositions in French and Italian political communication
di: Garassino, Davide, et al.
Pubblicazione: (2024)
di: Garassino, Davide, et al.
Pubblicazione: (2024)
The Silicon Ceiling: Auditing GPT's Race and Gender Biases in Hiring
di: Armstrong, Lena, et al.
Pubblicazione: (2024)
di: Armstrong, Lena, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Operationalizing Automated Essay Scoring: A Human-Aware Approach
di: Plasencia-Calaña, Yenisel
Pubblicazione: (2025) -
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
di: Chakravarty, Abhirup
Pubblicazione: (2025) -
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
di: Mohammadkhani, Ali Ghiasvand
Pubblicazione: (2024) -
Using GPT-4 to Augment Unbalanced Data for Automatic Scoring
di: Fang, Luyang, et al.
Pubblicazione: (2023) -
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals
di: Ehara, Yo
Pubblicazione: (2026)