Responsible AI for Test Equity and Quality: The Duolingo English Test as a Case Study
Fuente:
arXiv
Saved in:
| Main Authors: | Burstein, Jill, LaFlair, Geoffrey T., Yancey, Kevin, von Davier, Alina A., Dotan, Ravit |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Where Assessment Validation and Responsible AI Meet
by: Burstein, Jill, et al.
Published: (2024)
by: Burstein, Jill, et al.
Published: (2024)
BanditCAT and AutoIRT: Machine Learning Approaches to Computerized Adaptive Testing and Item Calibration
by: Sharpnack, James, et al.
Published: (2024)
by: Sharpnack, James, et al.
Published: (2024)
Developing an Automatic Pronunciation Scorer: Aligning Speech Evaluation Models and Applied Linguistics Constructs
by: Danwei Cai, et al.
Published: (2025)
by: Danwei Cai, et al.
Published: (2025)
Responsible Adoption of Generative AI in Higher Education: Developing a "Points to Consider" Approach Based on Faculty Perspectives
by: Dotan, Ravit, et al.
Published: (2024)
by: Dotan, Ravit, et al.
Published: (2024)
Exploring AI-Enabled Test Practice, Affect, and Test Outcomes in Language Assessment
by: Burstein, Jill, et al.
Published: (2025)
by: Burstein, Jill, et al.
Published: (2025)
AutoIRT: Calibrating Item Response Theory Models with Automated Machine Learning
by: Sharpnack, James, et al.
Published: (2024)
by: Sharpnack, James, et al.
Published: (2024)
Evolving AI Risk Management: A Maturity Model based on the NIST AI Risk Management Framework
by: Dotan, Ravit, et al.
Published: (2024)
by: Dotan, Ravit, et al.
Published: (2024)
Test Case-Informed Knowledge Tracing for Open-ended Coding Tasks
by: Duan, Zhangqi, et al.
Published: (2024)
by: Duan, Zhangqi, et al.
Published: (2024)
The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
by: Abdelnabi, Sahar, et al.
Published: (2025)
by: Abdelnabi, Sahar, et al.
Published: (2025)
AI Ethics by Design: Implementing Customizable Guardrails for Responsible AI Development
by: Šekrst, Kristina, et al.
Published: (2024)
by: Šekrst, Kristina, et al.
Published: (2024)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
by: Oh, Gyutaek, et al.
Published: (2025)
by: Oh, Gyutaek, et al.
Published: (2025)
Does Scientific Writing Converge to U.S. English? Evidence from Generative AI-Assisted Publications
by: Filimonovic, Dragan, et al.
Published: (2025)
by: Filimonovic, Dragan, et al.
Published: (2025)
Artificially Fluent: Swahili AI Performance Benchmarks Between English-Trained and Natively-Trained Datasets
by: Jaffer, Sophie, et al.
Published: (2025)
by: Jaffer, Sophie, et al.
Published: (2025)
Automated Question Generation for Science Tests in Arabic Language Using NLP Techniques
by: Tami, Mohammad, et al.
Published: (2024)
by: Tami, Mohammad, et al.
Published: (2024)
Legacy Procurement Practices Shape How U.S. Cities Govern AI: Understanding Government Employees' Practices, Challenges, and Needs
by: Johnson, Nari, et al.
Published: (2024)
by: Johnson, Nari, et al.
Published: (2024)
Developing Story: Case Studies of Generative AI's Use in Journalism
by: Brigham, Natalie Grace, et al.
Published: (2024)
by: Brigham, Natalie Grace, et al.
Published: (2024)
Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident
by: Borchers, Conrad, et al.
Published: (2026)
by: Borchers, Conrad, et al.
Published: (2026)
The Cost of Perfect English: Pragmatic Flattening and the Erasure of Authorial Voice in L2 Writing Supported by GenAI
by: Liu, Ao, et al.
Published: (2026)
by: Liu, Ao, et al.
Published: (2026)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024)
by: Jin, Yiping, et al.
Published: (2024)
Threefold model for AI Readiness: A Case Study with Finnish Healthcare SMEs
by: Alnajjar, Mohammed, et al.
Published: (2025)
by: Alnajjar, Mohammed, et al.
Published: (2025)
Passing the Turing Test in Political Discourse: Fine-Tuning LLMs to Mimic Polarized Social Media Comments
by: Pazzaglia, ., et al.
Published: (2025)
by: Pazzaglia, ., et al.
Published: (2025)
Testing network clustering algorithms with Natural Language Processing
by: Achitouv, Ixandra, et al.
Published: (2024)
by: Achitouv, Ixandra, et al.
Published: (2024)
Can Grammarly and ChatGPT accelerate language change? AI-powered technologies and their impact on the English language: wordiness vs. conciseness
by: Rudnicka, Karolina
Published: (2025)
by: Rudnicka, Karolina
Published: (2025)
Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
by: Li, Ming, et al.
Published: (2026)
by: Li, Ming, et al.
Published: (2026)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
by: Chehbouni, Khaoula, et al.
Published: (2024)
by: Chehbouni, Khaoula, et al.
Published: (2024)
Reinforcing Stereotypes of Anger: Emotion AI on African American Vernacular English
by: Dorn, Rebecca, et al.
Published: (2025)
by: Dorn, Rebecca, et al.
Published: (2025)
Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study
by: Silva, Jhessica, et al.
Published: (2025)
by: Silva, Jhessica, et al.
Published: (2025)
AGGA: A Dataset of Academic Guidelines for Generative AI and Large Language Models
by: Jiao, Junfeng, et al.
Published: (2025)
by: Jiao, Junfeng, et al.
Published: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course
by: Woo, David James, et al.
Published: (2026)
by: Woo, David James, et al.
Published: (2026)
Neutralizing the Narrative: AI-Powered Debiasing of Online News Articles
by: Kuo, Chen Wei, et al.
Published: (2025)
by: Kuo, Chen Wei, et al.
Published: (2025)
Gender Bias Detection in Court Decisions: A Brazilian Case Study
by: Benatti, Raysa, et al.
Published: (2024)
by: Benatti, Raysa, et al.
Published: (2024)
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions
by: Morabito, Robert, et al.
Published: (2024)
by: Morabito, Robert, et al.
Published: (2024)
Attributions toward Artificial Agents in a modified Moral Turing Test
by: Aharoni, Eyal, et al.
Published: (2024)
by: Aharoni, Eyal, et al.
Published: (2024)
Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings
by: Hong, Harbin, et al.
Published: (2025)
by: Hong, Harbin, et al.
Published: (2025)
Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation
by: Shen, Hanwen, et al.
Published: (2026)
by: Shen, Hanwen, et al.
Published: (2026)
Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study
by: Yaacoub, Antoun, et al.
Published: (2025)
by: Yaacoub, Antoun, et al.
Published: (2025)
Is Contrasting All You Need? Contrastive Learning for the Detection and Attribution of AI-generated Text
by: La Cava, Lucio, et al.
Published: (2024)
by: La Cava, Lucio, et al.
Published: (2024)
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
Automatic Speech Recognition Biases in Newcastle English: an Error Analysis
by: Serditova, Dana, et al.
Published: (2025)
by: Serditova, Dana, et al.
Published: (2025)
Similar Items
-
Where Assessment Validation and Responsible AI Meet
by: Burstein, Jill, et al.
Published: (2024) -
BanditCAT and AutoIRT: Machine Learning Approaches to Computerized Adaptive Testing and Item Calibration
by: Sharpnack, James, et al.
Published: (2024) -
Developing an Automatic Pronunciation Scorer: Aligning Speech Evaluation Models and Applied Linguistics Constructs
by: Danwei Cai, et al.
Published: (2025) -
Responsible Adoption of Generative AI in Higher Education: Developing a "Points to Consider" Approach Based on Faculty Perspectives
by: Dotan, Ravit, et al.
Published: (2024) -
Exploring AI-Enabled Test Practice, Affect, and Test Outcomes in Language Assessment
by: Burstein, Jill, et al.
Published: (2025)