Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Yishan, Borchers, Conrad, Cukurova, Mutlu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI for Accessible Education: Personalized Audio-Based Learning for Blind Students
by: Yang, Crystal, et al.
Published: (2025)
by: Yang, Crystal, et al.
Published: (2025)
LLMs as Educational Analysts: Transforming Multimodal Data Traces into Actionable Reading Assessment Reports
by: Davalos, Eduardo, et al.
Published: (2025)
by: Davalos, Eduardo, et al.
Published: (2025)
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
by: Alberts, Lize, et al.
Published: (2024)
by: Alberts, Lize, et al.
Published: (2024)
Collective Constitutional AI: Aligning a Language Model with Public Input
by: Huang, Saffron, et al.
Published: (2024)
by: Huang, Saffron, et al.
Published: (2024)
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
by: Mehta, Manisha, et al.
Published: (2025)
by: Mehta, Manisha, et al.
Published: (2025)
Aurora: Neuro-Symbolic AI Driven Advising Agent
by: Lugones, Lorena Amanda Quincoso, et al.
Published: (2026)
by: Lugones, Lorena Amanda Quincoso, et al.
Published: (2026)
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences
by: Baradari, Dünya, et al.
Published: (2025)
by: Baradari, Dünya, et al.
Published: (2025)
"I followed what felt right, not what I was told": Autonomy, Coaching, and Recognizing Bias Through AI-Mediated Dialogue
by: Taheri, Atieh, et al.
Published: (2026)
by: Taheri, Atieh, et al.
Published: (2026)
Developing Critical Thinking in Second Language Learners: Exploring Generative AI like ChatGPT as a Tool for Argumentative Essay Writing
by: Suh, Simon, et al.
Published: (2025)
by: Suh, Simon, et al.
Published: (2025)
Reviewriter: AI-Generated Instructions For Peer Review Writing
by: Su, Xiaotian, et al.
Published: (2025)
by: Su, Xiaotian, et al.
Published: (2025)
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
by: Zahid, Anwar Hossain, et al.
Published: (2025)
by: Zahid, Anwar Hossain, et al.
Published: (2025)
AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion Chatbot
by: Mohammad, et al.
Published: (2025)
by: Mohammad, et al.
Published: (2025)
How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses
by: Gupta, Ishan, et al.
Published: (2026)
by: Gupta, Ishan, et al.
Published: (2026)
How Adding Metacognitive Requirements in Support of AI Feedback in Practice Exams Transforms Student Learning Behaviors
by: Ahmad, Mak, et al.
Published: (2025)
by: Ahmad, Mak, et al.
Published: (2025)
Integrating Generative AI in Cybersecurity Education: Case Study Insights on Pedagogical Strategies, Critical Thinking, and Responsible AI Use
by: Elkhodr, Mahmoud, et al.
Published: (2025)
by: Elkhodr, Mahmoud, et al.
Published: (2025)
Human-in-the-Loop Benchmarking of Heterogeneous LLMs for Automated Competency Assessment in Secondary Level Mathematics
by: Bhusal, Jatin, et al.
Published: (2026)
by: Bhusal, Jatin, et al.
Published: (2026)
Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment
by: Burleigh, Tyler
Published: (2026)
by: Burleigh, Tyler
Published: (2026)
Generative UI as an Accessibility Bridge: Lessons from C2C E-Commerce
by: Ryskeldiev, Bektur
Published: (2026)
by: Ryskeldiev, Bektur
Published: (2026)
Cognitive Twins: Investigating Personalized Thinking Model Building and Its Performance Enhancement with Human-in-the-Loop
by: Hwang, Wu-Yuin, et al.
Published: (2026)
by: Hwang, Wu-Yuin, et al.
Published: (2026)
Textual Entailment is not a Better Bias Metric than Token Probability
by: Felkner, Virginia K., et al.
Published: (2025)
by: Felkner, Virginia K., et al.
Published: (2025)
The Human-AI Delegation Dilemma: Individual Strategies, Collective Equilibria and Sociotechnical Lock-in
by: Hila, Angjelin
Published: (2026)
by: Hila, Angjelin
Published: (2026)
The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims
by: Kelly, Matthew
Published: (2025)
by: Kelly, Matthew
Published: (2025)
How AI Systems Think About Education: Analyzing Latent Preference Patterns in Large Language Models
by: Autenrieth, Daniel
Published: (2026)
by: Autenrieth, Daniel
Published: (2026)
Detecting AI-Assisted Cheating in Online Exams through Behavior Analytics
by: Akçapınar, Gökhan
Published: (2025)
by: Akçapınar, Gökhan
Published: (2025)
The Reliance Negotiation Framework: A Dynamic Process Model of Student LLM Engagement in Academic Writing
by: Hossain, Shahin
Published: (2026)
by: Hossain, Shahin
Published: (2026)
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
by: Felkner, Virginia K., et al.
Published: (2024)
by: Felkner, Virginia K., et al.
Published: (2024)
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
by: Maheshwari, Aditya, et al.
Published: (2026)
by: Maheshwari, Aditya, et al.
Published: (2026)
Chatbot Deployment Considerations for Application-Agnostic Human-Machine Dialogues
by: Rivas, Pablo, et al.
Published: (2025)
by: Rivas, Pablo, et al.
Published: (2025)
Balancing Innovation and Integrity: AI Integration in Liberal Arts College Administration
by: Read, Ian Olivo
Published: (2025)
by: Read, Ian Olivo
Published: (2025)
Extreme Self-Preference in Language Models
by: Lehr, Steven A., et al.
Published: (2025)
by: Lehr, Steven A., et al.
Published: (2025)
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
by: Kuric, Eduard, et al.
Published: (2026)
by: Kuric, Eduard, et al.
Published: (2026)
Agentic Education: Using Claude Code to Teach Claude Code
by: Naboulsi, Zain
Published: (2026)
by: Naboulsi, Zain
Published: (2026)
Platform-Independent and Curriculum-Oriented Intelligent Assistant for Higher Education
by: Sajja, Ramteja, et al.
Published: (2023)
by: Sajja, Ramteja, et al.
Published: (2023)
CS-Guide: Leveraging LLMs and Student Reflections to Provide Frequent, Scalable Academic Monitoring Feedback to Computer Science Students
by: Chacko, Samuel Jacob, et al.
Published: (2025)
by: Chacko, Samuel Jacob, et al.
Published: (2025)
Using Sentiment Analysis to Investigate Peer Feedback by Native and Non-Native English Speakers
by: Exline, Brittney, et al.
Published: (2025)
by: Exline, Brittney, et al.
Published: (2025)
Can Humans Tell? A Dual-Axis Study of Human Perception of LLM-Generated News
by: Loth, Alexander, et al.
Published: (2026)
by: Loth, Alexander, et al.
Published: (2026)
A Contextual Help Browser Extension to Assist Digital Illiterate Internet Users
by: Koutsiaris, Christos
Published: (2026)
by: Koutsiaris, Christos
Published: (2026)
Evaluating Epistemic Guardrails in AI Reading Assistants: A Behavioral Audit of a Minimal Prototype
by: Agustin, Matthew Christian
Published: (2026)
by: Agustin, Matthew Christian
Published: (2026)
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
by: Sami, K. M. Jubair, et al.
Published: (2026)
by: Sami, K. M. Jubair, et al.
Published: (2026)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
by: Hari, Vishnu, et al.
Published: (2025)
by: Hari, Vishnu, et al.
Published: (2025)
Similar Items
-
AI for Accessible Education: Personalized Audio-Based Learning for Blind Students
by: Yang, Crystal, et al.
Published: (2025) -
LLMs as Educational Analysts: Transforming Multimodal Data Traces into Actionable Reading Assessment Reports
by: Davalos, Eduardo, et al.
Published: (2025) -
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
by: Alberts, Lize, et al.
Published: (2024) -
Collective Constitutional AI: Aligning a Language Model with Public Input
by: Huang, Saffron, et al.
Published: (2024) -
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
by: Mehta, Manisha, et al.
Published: (2025)