Implicit Grading Bias in Large Language Models: How Writing Style Affects Automated Assessment Across Math, Programming, and Essay Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Jadhav, Rudra, Danve, Janhavi, Shaw, Sonalika |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era
by: Jadhav, Rudra, et al.
Published: (2026)
by: Jadhav, Rudra, et al.
Published: (2026)
Using Large Language Models for Automated Grading of Student Writing about Science
by: Impey, Chris, et al.
Published: (2024)
by: Impey, Chris, et al.
Published: (2024)
Empirical Study of Large Language Models as Automated Essay Scoring Tools in English Composition__Taking TOEFL Independent Writing Task for Example
by: Xia, Wei, et al.
Published: (2024)
by: Xia, Wei, et al.
Published: (2024)
Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
by: Hao, Jiangang
Published: (2026)
by: Hao, Jiangang
Published: (2026)
EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
by: Gao, Fan, et al.
Published: (2025)
by: Gao, Fan, et al.
Published: (2025)
Benchmarking Large Language Models for Math Reasoning Tasks
by: Seßler, Kathrin, et al.
Published: (2024)
by: Seßler, Kathrin, et al.
Published: (2024)
Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
by: Xu, Qingshu, et al.
Published: (2025)
by: Xu, Qingshu, et al.
Published: (2025)
Evaluating Austrian A-Level German Essays with Large Language Models for Automated Essay Scoring
by: Kubesch, Jonas, et al.
Published: (2026)
by: Kubesch, Jonas, et al.
Published: (2026)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
by: Bai, Xuechunzi, et al.
Published: (2024)
by: Bai, Xuechunzi, et al.
Published: (2024)
How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
by: Jung, Julie, et al.
Published: (2025)
by: Jung, Julie, et al.
Published: (2025)
How well can LLMs Grade Essays in Arabic?
by: Ghazawi, Rayed, et al.
Published: (2025)
by: Ghazawi, Rayed, et al.
Published: (2025)
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
by: Tomar, Aditya, et al.
Published: (2025)
by: Tomar, Aditya, et al.
Published: (2025)
EssayCBM: Rubric-Aligned Concept Bottleneck Models for Transparent Essay Grading
by: Chaudhary, Kumar Satvik, et al.
Published: (2025)
by: Chaudhary, Kumar Satvik, et al.
Published: (2025)
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
by: Cai, Yida, et al.
Published: (2025)
by: Cai, Yida, et al.
Published: (2025)
IELTS Writing Revision Platform with Automated Essay Scoring and Adaptive Feedback
by: Ramancauskas, Titas, et al.
Published: (2025)
by: Ramancauskas, Titas, et al.
Published: (2025)
Long Context Automated Essay Scoring with Language Models
by: Ormerod, Christopher, et al.
Published: (2025)
by: Ormerod, Christopher, et al.
Published: (2025)
The Widespread Adoption of Large Language Model-Assisted Writing Across Society
by: Liang, Weixin, et al.
Published: (2025)
by: Liang, Weixin, et al.
Published: (2025)
Hey AI Can You Grade My Essay?: Automatic Essay Grading
by: Maliha, Maisha, et al.
Published: (2024)
by: Maliha, Maisha, et al.
Published: (2024)
Does the Prompt-based Large Language Model Recognize Students' Demographics and Introduce Bias in Essay Scoring?
by: Yang, Kaixun, et al.
Published: (2025)
by: Yang, Kaixun, et al.
Published: (2025)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
by: Su, Jiamin, et al.
Published: (2025)
by: Su, Jiamin, et al.
Published: (2025)
Toward Trustworthy Difficulty Assessments: Large Language Models as Judges in Programming and Synthetic Tasks
by: Tabib, H. M. Shadman, et al.
Published: (2025)
by: Tabib, H. M. Shadman, et al.
Published: (2025)
Learning to Write Rationally: How Information Is Distributed in Non-Native Speakers' Essays
by: Tang, Zixin, et al.
Published: (2024)
by: Tang, Zixin, et al.
Published: (2024)
How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations
by: Takenami, Yoshiki, et al.
Published: (2025)
by: Takenami, Yoshiki, et al.
Published: (2025)
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
by: Vedula, Bhaskara Hanuma, et al.
Published: (2026)
by: Vedula, Bhaskara Hanuma, et al.
Published: (2026)
Exploring Automated Distractor Generation for Math Multiple-choice Questions via Large Language Models
by: Feng, Wanyong, et al.
Published: (2024)
by: Feng, Wanyong, et al.
Published: (2024)
Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset
by: Wang, Leroy Z.
Published: (2025)
by: Wang, Leroy Z.
Published: (2025)
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
by: Chakravarty, Abhirup
Published: (2025)
by: Chakravarty, Abhirup
Published: (2025)
Red-Teaming for Inducing Societal Bias in Large Language Models
by: Luo, Chu Fei, et al.
Published: (2024)
by: Luo, Chu Fei, et al.
Published: (2024)
How Does Code Pretraining Affect Language Model Task Performance?
by: Petty, Jackson, et al.
Published: (2024)
by: Petty, Jackson, et al.
Published: (2024)
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
by: Ahuja, Sanchit, et al.
Published: (2023)
by: Ahuja, Sanchit, et al.
Published: (2023)
Inference-Time Reasoning Selectively Reduces Implicit Social Bias in Large Language Models
by: Apsel, Molly, et al.
Published: (2026)
by: Apsel, Molly, et al.
Published: (2026)
Do We Need a Detailed Rubric for Automated Essay Scoring using Large Language Models?
by: Yoshida, Lui
Published: (2025)
by: Yoshida, Lui
Published: (2025)
How Does the Disclosure of AI Assistance Affect the Perceptions of Writing?
by: Li, Zhuoyan, et al.
Published: (2024)
by: Li, Zhuoyan, et al.
Published: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
by: Ying, Huaiyuan, et al.
Published: (2024)
by: Ying, Huaiyuan, et al.
Published: (2024)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
by: Mitra, Arindam, et al.
Published: (2024)
by: Mitra, Arindam, et al.
Published: (2024)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
by: Truong, Kimberly Le, et al.
Published: (2025)
by: Truong, Kimberly Le, et al.
Published: (2025)
Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection
by: Zhao, Yachao, et al.
Published: (2025)
by: Zhao, Yachao, et al.
Published: (2025)
Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems
by: Ye, Tian, et al.
Published: (2024)
by: Ye, Tian, et al.
Published: (2024)
TaskBench: Benchmarking Large Language Models for Task Automation
by: Shen, Yongliang, et al.
Published: (2023)
by: Shen, Yongliang, et al.
Published: (2023)
Large Language Models Struggle with Unreasonability in Math Problems
by: Ma, Jingyuan, et al.
Published: (2024)
by: Ma, Jingyuan, et al.
Published: (2024)
Similar Items
-
The AI Skills Shift: Mapping Skill Obsolescence, Emergence, and Transition Pathways in the LLM Era
by: Jadhav, Rudra, et al.
Published: (2026) -
Using Large Language Models for Automated Grading of Student Writing about Science
by: Impey, Chris, et al.
Published: (2024) -
Empirical Study of Large Language Models as Automated Essay Scoring Tools in English Composition__Taking TOEFL Independent Writing Task for Example
by: Xia, Wei, et al.
Published: (2024) -
Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
by: Hao, Jiangang
Published: (2026) -
EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
by: Gao, Fan, et al.
Published: (2025)