Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach
Fuente:
arXiv
Salvato in:
| Autori principali: | Casabianca, Jodi M., Beiting-Parrish, Maggie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Scaling Law of Evaluation Failure: Why Simple Averaging Collapses Under Data Sparsity and Item Difficulty Gaps, and How Item Response Theory Recovers Ground Truth Across Domains
di: Kang, Jung Min
Pubblicazione: (2026)
di: Kang, Jung Min
Pubblicazione: (2026)
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
di: Quaye, Jessica, et al.
Pubblicazione: (2025)
di: Quaye, Jessica, et al.
Pubblicazione: (2025)
Can Active Label Correction Improve LLM-based Modular AI Systems?
di: Taneja, Karan, et al.
Pubblicazione: (2024)
di: Taneja, Karan, et al.
Pubblicazione: (2024)
MLV$^2$-Net: Rater-Based Majority-Label Voting for Consistent Meningeal Lymphatic Vessel Segmentation
di: Bongratz, Fabian, et al.
Pubblicazione: (2024)
di: Bongratz, Fabian, et al.
Pubblicazione: (2024)
DataRater: Meta-Learned Dataset Curation
di: Calian, Dan A., et al.
Pubblicazione: (2025)
di: Calian, Dan A., et al.
Pubblicazione: (2025)
Efficient Bilevel Optimization for Meta Label Correction in Noisy Label Learning
di: Nguyen, Ba Hoang Anh, et al.
Pubblicazione: (2026)
di: Nguyen, Ba Hoang Anh, et al.
Pubblicazione: (2026)
TMLC-Net: Transferable Meta Label Correction for Noisy Label Learning
di: Li, Mengyang
Pubblicazione: (2025)
di: Li, Mengyang
Pubblicazione: (2025)
Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach
di: Jurenka, Irina, et al.
Pubblicazione: (2024)
di: Jurenka, Irina, et al.
Pubblicazione: (2024)
Holistic Safety and Responsibility Evaluations of Advanced AI Models
di: Weidinger, Laura, et al.
Pubblicazione: (2024)
di: Weidinger, Laura, et al.
Pubblicazione: (2024)
Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence
di: Resnick, Paul, et al.
Pubblicazione: (2021)
di: Resnick, Paul, et al.
Pubblicazione: (2021)
Learning to Clean: Reinforcement Learning for Noisy Label Correction
di: Heidari, Marzi, et al.
Pubblicazione: (2025)
di: Heidari, Marzi, et al.
Pubblicazione: (2025)
Label Noise Robustness for Domain-Agnostic Fair Corrections via Nearest Neighbors Label Spreading
di: Stromberg, Nathan, et al.
Pubblicazione: (2024)
di: Stromberg, Nathan, et al.
Pubblicazione: (2024)
From Feature-Based Models to Generative AI: Validity Evidence for Constructed Response Scoring
di: Casabianca, Jodi M., et al.
Pubblicazione: (2026)
di: Casabianca, Jodi M., et al.
Pubblicazione: (2026)
GFLC: Graph-based Fairness-aware Label Correction for Fair Classification
di: Sulaiman, Modar, et al.
Pubblicazione: (2025)
di: Sulaiman, Modar, et al.
Pubblicazione: (2025)
Tackling Noisy Clients in Federated Learning with End-to-end Label Correction
di: Jiang, Xuefeng, et al.
Pubblicazione: (2024)
di: Jiang, Xuefeng, et al.
Pubblicazione: (2024)
Approaches to Responsible Governance of GenAI in Organizations
di: Gandhi, Dhari, et al.
Pubblicazione: (2025)
di: Gandhi, Dhari, et al.
Pubblicazione: (2025)
SAP: Corrective Machine Unlearning with Scaled Activation Projection for Label Noise Robustness
di: Kodge, Sangamesh, et al.
Pubblicazione: (2024)
di: Kodge, Sangamesh, et al.
Pubblicazione: (2024)
Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs
di: Dong, Zihan, et al.
Pubblicazione: (2026)
di: Dong, Zihan, et al.
Pubblicazione: (2026)
Correcting Noisy Multilabel Predictions: Modeling Label Noise through Latent Space Shifts
di: Huang, Weipeng, et al.
Pubblicazione: (2025)
di: Huang, Weipeng, et al.
Pubblicazione: (2025)
FedEFC: Federated Learning Using Enhanced Forward Correction Against Noisy Labels
di: Yu, Seunghun, et al.
Pubblicazione: (2025)
di: Yu, Seunghun, et al.
Pubblicazione: (2025)
AI Alignment via Incentives and Correction
di: Agarwal, Rohit, et al.
Pubblicazione: (2026)
di: Agarwal, Rohit, et al.
Pubblicazione: (2026)
NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training
di: Wu, Fang, et al.
Pubblicazione: (2026)
di: Wu, Fang, et al.
Pubblicazione: (2026)
Variance-Bounded Evaluation of Entity-Centric AI Systems Without Ground Truth: Theory and Measurement
di: Ding, Kaihua
Pubblicazione: (2025)
di: Ding, Kaihua
Pubblicazione: (2025)
pAI/MSc: ML Theory Research with Humans on the Loop
di: Abdelmoneum, Mahmoud, et al.
Pubblicazione: (2026)
di: Abdelmoneum, Mahmoud, et al.
Pubblicazione: (2026)
LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems
di: Li, Chu, et al.
Pubblicazione: (2024)
di: Li, Chu, et al.
Pubblicazione: (2024)
A Confidence-Variance Theory for Pseudo-Label Selection in Semi-Supervised Learning
di: Liu, Jinshi, et al.
Pubblicazione: (2026)
di: Liu, Jinshi, et al.
Pubblicazione: (2026)
Who's Gaming the System? A Causally-Motivated Approach for Detecting Strategic Adaptation
di: Chang, Trenton, et al.
Pubblicazione: (2024)
di: Chang, Trenton, et al.
Pubblicazione: (2024)
Predicting Effects, Missing Distributions: Evaluating LLMs as Human Behavior Simulators in Operations Management
di: Zhang, Runze, et al.
Pubblicazione: (2025)
di: Zhang, Runze, et al.
Pubblicazione: (2025)
Revisiting Rogers' Paradox in the Context of Human-AI Interaction
di: Collins, Katherine M., et al.
Pubblicazione: (2025)
di: Collins, Katherine M., et al.
Pubblicazione: (2025)
NudgeRank: Digital Algorithmic Nudging for Personalized Health
di: Chiam, Jodi, et al.
Pubblicazione: (2024)
di: Chiam, Jodi, et al.
Pubblicazione: (2024)
Towards a Theory of AI Personhood
di: Ward, Francis Rhys
Pubblicazione: (2025)
di: Ward, Francis Rhys
Pubblicazione: (2025)
Position: AI Evaluations Should be Grounded on a Theory of Capability
di: Jo, Nathanael, et al.
Pubblicazione: (2025)
di: Jo, Nathanael, et al.
Pubblicazione: (2025)
Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs
di: Flores, Gerardo A., et al.
Pubblicazione: (2025)
di: Flores, Gerardo A., et al.
Pubblicazione: (2025)
Monodense Deep Neural Model for Determining Item Price Elasticity
di: Garg, Lakshya, et al.
Pubblicazione: (2026)
di: Garg, Lakshya, et al.
Pubblicazione: (2026)
A Masked Semi-Supervised Learning Approach for Otago Micro Labels Recognition
di: Shang, Meng, et al.
Pubblicazione: (2024)
di: Shang, Meng, et al.
Pubblicazione: (2024)
On the Error-Correcting Effects of Stochasticity in Discrete Diffusion
di: Yuan, William, et al.
Pubblicazione: (2026)
di: Yuan, William, et al.
Pubblicazione: (2026)
Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length
di: Xu, Zhiyu, et al.
Pubblicazione: (2025)
di: Xu, Zhiyu, et al.
Pubblicazione: (2025)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
di: Rogoz, Ana-Cristina, et al.
Pubblicazione: (2024)
di: Rogoz, Ana-Cristina, et al.
Pubblicazione: (2024)
SAFE: A Novel Approach to AI Weather Evaluation through Stratified Assessments of Forecasts over Earth
di: Masi, Nick, et al.
Pubblicazione: (2025)
di: Masi, Nick, et al.
Pubblicazione: (2025)
A Descriptive and Normative Theory of Human Beliefs in RLHF
di: Dandekar, Sylee, et al.
Pubblicazione: (2025)
di: Dandekar, Sylee, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Scaling Law of Evaluation Failure: Why Simple Averaging Collapses Under Data Sparsity and Item Difficulty Gaps, and How Item Response Theory Recovers Ground Truth Across Domains
di: Kang, Jung Min
Pubblicazione: (2026) -
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
di: Quaye, Jessica, et al.
Pubblicazione: (2025) -
Can Active Label Correction Improve LLM-based Modular AI Systems?
di: Taneja, Karan, et al.
Pubblicazione: (2024) -
MLV$^2$-Net: Rater-Based Majority-Label Voting for Consistent Meningeal Lymphatic Vessel Segmentation
di: Bongratz, Fabian, et al.
Pubblicazione: (2024) -
DataRater: Meta-Learned Dataset Curation
di: Calian, Dan A., et al.
Pubblicazione: (2025)