Implicit Grading Bias in Large Language Models: How Writing Style Affects Automated Assessment Across Math, Programming, and Essay Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Jadhav, Rudra, Danve, Janhavi, Shaw, Sonalika |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era
di: Jadhav, Rudra, et al.
Pubblicazione: (2026)
di: Jadhav, Rudra, et al.
Pubblicazione: (2026)
Using Large Language Models for Automated Grading of Student Writing about Science
di: Impey, Chris, et al.
Pubblicazione: (2024)
di: Impey, Chris, et al.
Pubblicazione: (2024)
Empirical Study of Large Language Models as Automated Essay Scoring Tools in English Composition__Taking TOEFL Independent Writing Task for Example
di: Xia, Wei, et al.
Pubblicazione: (2024)
di: Xia, Wei, et al.
Pubblicazione: (2024)
Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
di: Hao, Jiangang
Pubblicazione: (2026)
di: Hao, Jiangang
Pubblicazione: (2026)
EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
di: Gao, Fan, et al.
Pubblicazione: (2025)
di: Gao, Fan, et al.
Pubblicazione: (2025)
Benchmarking Large Language Models for Math Reasoning Tasks
di: Seßler, Kathrin, et al.
Pubblicazione: (2024)
di: Seßler, Kathrin, et al.
Pubblicazione: (2024)
Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
di: Xu, Qingshu, et al.
Pubblicazione: (2025)
di: Xu, Qingshu, et al.
Pubblicazione: (2025)
Evaluating Austrian A-Level German Essays with Large Language Models for Automated Essay Scoring
di: Kubesch, Jonas, et al.
Pubblicazione: (2026)
di: Kubesch, Jonas, et al.
Pubblicazione: (2026)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
di: Bai, Xuechunzi, et al.
Pubblicazione: (2024)
di: Bai, Xuechunzi, et al.
Pubblicazione: (2024)
How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
di: Jung, Julie, et al.
Pubblicazione: (2025)
di: Jung, Julie, et al.
Pubblicazione: (2025)
How well can LLMs Grade Essays in Arabic?
di: Ghazawi, Rayed, et al.
Pubblicazione: (2025)
di: Ghazawi, Rayed, et al.
Pubblicazione: (2025)
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
di: Tomar, Aditya, et al.
Pubblicazione: (2025)
di: Tomar, Aditya, et al.
Pubblicazione: (2025)
EssayCBM: Rubric-Aligned Concept Bottleneck Models for Transparent Essay Grading
di: Chaudhary, Kumar Satvik, et al.
Pubblicazione: (2025)
di: Chaudhary, Kumar Satvik, et al.
Pubblicazione: (2025)
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
di: Cai, Yida, et al.
Pubblicazione: (2025)
di: Cai, Yida, et al.
Pubblicazione: (2025)
IELTS Writing Revision Platform with Automated Essay Scoring and Adaptive Feedback
di: Ramancauskas, Titas, et al.
Pubblicazione: (2025)
di: Ramancauskas, Titas, et al.
Pubblicazione: (2025)
Long Context Automated Essay Scoring with Language Models
di: Ormerod, Christopher, et al.
Pubblicazione: (2025)
di: Ormerod, Christopher, et al.
Pubblicazione: (2025)
The Widespread Adoption of Large Language Model-Assisted Writing Across Society
di: Liang, Weixin, et al.
Pubblicazione: (2025)
di: Liang, Weixin, et al.
Pubblicazione: (2025)
Hey AI Can You Grade My Essay?: Automatic Essay Grading
di: Maliha, Maisha, et al.
Pubblicazione: (2024)
di: Maliha, Maisha, et al.
Pubblicazione: (2024)
Does the Prompt-based Large Language Model Recognize Students' Demographics and Introduce Bias in Essay Scoring?
di: Yang, Kaixun, et al.
Pubblicazione: (2025)
di: Yang, Kaixun, et al.
Pubblicazione: (2025)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
di: Su, Jiamin, et al.
Pubblicazione: (2025)
di: Su, Jiamin, et al.
Pubblicazione: (2025)
Toward Trustworthy Difficulty Assessments: Large Language Models as Judges in Programming and Synthetic Tasks
di: Tabib, H. M. Shadman, et al.
Pubblicazione: (2025)
di: Tabib, H. M. Shadman, et al.
Pubblicazione: (2025)
Learning to Write Rationally: How Information Is Distributed in Non-Native Speakers' Essays
di: Tang, Zixin, et al.
Pubblicazione: (2024)
di: Tang, Zixin, et al.
Pubblicazione: (2024)
How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations
di: Takenami, Yoshiki, et al.
Pubblicazione: (2025)
di: Takenami, Yoshiki, et al.
Pubblicazione: (2025)
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
di: Vedula, Bhaskara Hanuma, et al.
Pubblicazione: (2026)
di: Vedula, Bhaskara Hanuma, et al.
Pubblicazione: (2026)
Exploring Automated Distractor Generation for Math Multiple-choice Questions via Large Language Models
di: Feng, Wanyong, et al.
Pubblicazione: (2024)
di: Feng, Wanyong, et al.
Pubblicazione: (2024)
Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset
di: Wang, Leroy Z.
Pubblicazione: (2025)
di: Wang, Leroy Z.
Pubblicazione: (2025)
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
di: Chakravarty, Abhirup
Pubblicazione: (2025)
di: Chakravarty, Abhirup
Pubblicazione: (2025)
Red-Teaming for Inducing Societal Bias in Large Language Models
di: Luo, Chu Fei, et al.
Pubblicazione: (2024)
di: Luo, Chu Fei, et al.
Pubblicazione: (2024)
How Does Code Pretraining Affect Language Model Task Performance?
di: Petty, Jackson, et al.
Pubblicazione: (2024)
di: Petty, Jackson, et al.
Pubblicazione: (2024)
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
di: Ahuja, Sanchit, et al.
Pubblicazione: (2023)
di: Ahuja, Sanchit, et al.
Pubblicazione: (2023)
Inference-Time Reasoning Selectively Reduces Implicit Social Bias in Large Language Models
di: Apsel, Molly, et al.
Pubblicazione: (2026)
di: Apsel, Molly, et al.
Pubblicazione: (2026)
Do We Need a Detailed Rubric for Automated Essay Scoring using Large Language Models?
di: Yoshida, Lui
Pubblicazione: (2025)
di: Yoshida, Lui
Pubblicazione: (2025)
How Does the Disclosure of AI Assistance Affect the Perceptions of Writing?
di: Li, Zhuoyan, et al.
Pubblicazione: (2024)
di: Li, Zhuoyan, et al.
Pubblicazione: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
di: Ying, Huaiyuan, et al.
Pubblicazione: (2024)
di: Ying, Huaiyuan, et al.
Pubblicazione: (2024)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
di: Mitra, Arindam, et al.
Pubblicazione: (2024)
di: Mitra, Arindam, et al.
Pubblicazione: (2024)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
di: Truong, Kimberly Le, et al.
Pubblicazione: (2025)
di: Truong, Kimberly Le, et al.
Pubblicazione: (2025)
Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection
di: Zhao, Yachao, et al.
Pubblicazione: (2025)
di: Zhao, Yachao, et al.
Pubblicazione: (2025)
Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems
di: Ye, Tian, et al.
Pubblicazione: (2024)
di: Ye, Tian, et al.
Pubblicazione: (2024)
TaskBench: Benchmarking Large Language Models for Task Automation
di: Shen, Yongliang, et al.
Pubblicazione: (2023)
di: Shen, Yongliang, et al.
Pubblicazione: (2023)
Large Language Models Struggle with Unreasonability in Math Problems
di: Ma, Jingyuan, et al.
Pubblicazione: (2024)
di: Ma, Jingyuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era
di: Jadhav, Rudra, et al.
Pubblicazione: (2026) -
Using Large Language Models for Automated Grading of Student Writing about Science
di: Impey, Chris, et al.
Pubblicazione: (2024) -
Empirical Study of Large Language Models as Automated Essay Scoring Tools in English Composition__Taking TOEFL Independent Writing Task for Example
di: Xia, Wei, et al.
Pubblicazione: (2024) -
Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
di: Hao, Jiangang
Pubblicazione: (2026) -
EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
di: Gao, Fan, et al.
Pubblicazione: (2025)