Validating LLM-as-a-Judge Systems under Rating Indeterminacy
Fuente:
arXiv
Saved in:
| Main Authors: | Guerdan, Luke, Barocas, Solon, Holstein, Kenneth, Wallach, Hanna, Wu, Zhiwei Steven, Chouldechova, Alexandra |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024)
by: Guerdan, Luke, et al.
Published: (2024)
Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks
by: Guerdan, Luke, et al.
Published: (2025)
by: Guerdan, Luke, et al.
Published: (2025)
Effects of Generative AI Errors on User Reliance Across Task Difficulty
by: Anthis, Jacy Reese, et al.
Published: (2026)
by: Anthis, Jacy Reese, et al.
Published: (2026)
Do Responsible AI Artifacts Advance Stakeholder Goals? Four Key Barriers Perceived by Legal and Civil Stakeholders
by: Kawakami, Anna, et al.
Published: (2024)
by: Kawakami, Anna, et al.
Published: (2024)
Predictive Performance Comparison of Decision Policies Under Confounding
by: Guerdan, Luke, et al.
Published: (2024)
by: Guerdan, Luke, et al.
Published: (2024)
Creative Loss: Ambiguity, Uncertainty and Indeterminacy
by: Holberton, Tom
Published: (2024)
by: Holberton, Tom
Published: (2024)
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
by: Pang, Rock Yuren, et al.
Published: (2025)
by: Pang, Rock Yuren, et al.
Published: (2025)
Funding AI for Good: A Call for Meaningful Engagement
by: Lin, Hongjin, et al.
Published: (2025)
by: Lin, Hongjin, et al.
Published: (2025)
Dimensions of Generative AI Evaluation Design
by: Dow, P. Alex, et al.
Published: (2024)
by: Dow, P. Alex, et al.
Published: (2024)
Supporting Industry Computing Researchers in Assessing, Articulating, and Addressing the Potential Negative Societal Impact of Their Work
by: Deng, Wesley Hanwen, et al.
Published: (2024)
by: Deng, Wesley Hanwen, et al.
Published: (2024)
Leveraging Expert Consistency to Improve Algorithmic Decision Support
by: De-Arteaga, Maria, et al.
Published: (2021)
by: De-Arteaga, Maria, et al.
Published: (2021)
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
by: Lucy, Li, et al.
Published: (2023)
by: Lucy, Li, et al.
Published: (2023)
Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
by: Guerdan, Luke, et al.
Published: (2025)
by: Guerdan, Luke, et al.
Published: (2025)
VayuBuddy: an LLM-Powered Chatbot to Democratize Air Quality Insights
by: Patel, Zeel B, et al.
Published: (2024)
by: Patel, Zeel B, et al.
Published: (2024)
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026)
by: Agarwal, Dhruv, et al.
Published: (2026)
"Which LLM should I use?": Evaluating LLMs for tasks performed by Undergraduate Computer Science Students
by: Agarwal, Vibhor, et al.
Published: (2024)
by: Agarwal, Vibhor, et al.
Published: (2024)
Judging the algorithm: Algorithmic accountability on the risk assessment tool for intimate partner violence in the Basque Country
by: Valdivia, Ana, et al.
Published: (2022)
by: Valdivia, Ana, et al.
Published: (2022)
iLLuMinaTE: An LLM-XAI Framework Leveraging Social Science Explanation Theories Towards Actionable Student Performance Feedback
by: Swamy, Vinitra, et al.
Published: (2024)
by: Swamy, Vinitra, et al.
Published: (2024)
Driving Assistance System for Ambulances to Minimise the Vibrations in Patient Cabin
by: Aldegheishem, Abdulaziz, et al.
Published: (2026)
by: Aldegheishem, Abdulaziz, et al.
Published: (2026)
The Trust Calibration Maturity Model for Characterizing and Communicating Trustworthiness of AI Systems
by: Steinmetz, Scott T, et al.
Published: (2025)
by: Steinmetz, Scott T, et al.
Published: (2025)
Wisdom of the LLM Crowd: A Large Scale Benchmark of Multi-Label U.S. Election-Related Harmful Social Media Content
by: Wang, Qile, et al.
Published: (2026)
by: Wang, Qile, et al.
Published: (2026)
PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
by: Deng, Wesley Hanwen, et al.
Published: (2026)
by: Deng, Wesley Hanwen, et al.
Published: (2026)
Platforms as Crime Scene, Judge, and Jury: How Victim-Survivors of Non-Consensual Intimate Imagery Report Abuse Online
by: Qiwei, Li, et al.
Published: (2025)
by: Qiwei, Li, et al.
Published: (2025)
A Nested Model for AI Design and Validation
by: Dubey, Akshat, et al.
Published: (2024)
by: Dubey, Akshat, et al.
Published: (2024)
Enhancing LLM-Based Feedback: Insights from Intelligent Tutoring Systems and the Learning Sciences
by: Stamper, John, et al.
Published: (2024)
by: Stamper, John, et al.
Published: (2024)
Low-Cost System for Automatic Recognition of Driving Pattern in Assessing Interurban Mobility using Geo-Information
by: Romero, Oscar, et al.
Published: (2026)
by: Romero, Oscar, et al.
Published: (2026)
Towards better social crisis data with HERMES: Hybrid sensing for EmeRgency ManagEment System
by: Avvenuti, Marco, et al.
Published: (2019)
by: Avvenuti, Marco, et al.
Published: (2019)
A Fast and Minimal System to Identify Depression Using Smartphones: Explainable Machine Learning-Based Approach
by: Ahmed, Md Sabbir, et al.
Published: (2025)
by: Ahmed, Md Sabbir, et al.
Published: (2025)
ChatISA: A Prompt-Engineered, In-House Multi-Modal Generative AI Chatbot for Information Systems Education
by: Megahed, Fadel M., et al.
Published: (2024)
by: Megahed, Fadel M., et al.
Published: (2024)
A Taxonomy of Questions for Critical Reflection in Machine-Assisted Decision-Making
by: Fischer, Simon W. S., et al.
Published: (2025)
by: Fischer, Simon W. S., et al.
Published: (2025)
LLM-based Cognitive Models of Students with Misconceptions
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
The Role of Visualization in LLM-Assisted Knowledge Graph Systems: Effects on User Trust, Exploration, and Workflows
by: Li, Harry, et al.
Published: (2025)
by: Li, Harry, et al.
Published: (2025)
Cross-Cultural Validation of Partner Models for Voice User Interfaces
by: Seaborn, Katie, et al.
Published: (2024)
by: Seaborn, Katie, et al.
Published: (2024)
Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz
by: Gmeiner, Frederic, et al.
Published: (2025)
by: Gmeiner, Frederic, et al.
Published: (2025)
CodeTailor: LLM-Powered Personalized Parsons Puzzles for Engaging Support While Learning Programming
by: Hou, Xinying, et al.
Published: (2024)
by: Hou, Xinying, et al.
Published: (2024)
Constraining Participation: Affordances of Feedback Features in Interfaces to Large Language Models
by: Cooper, Ned, et al.
Published: (2024)
by: Cooper, Ned, et al.
Published: (2024)
Personalized Parsons Puzzles as Scaffolding Enhance Practice Engagement Over Just Showing LLM-Powered Solutions
by: Hou, Xinying, et al.
Published: (2025)
by: Hou, Xinying, et al.
Published: (2025)
Adversarial Attacks and Defenses in Physiological Computing: A Systematic Review
by: Wu, Dongrui, et al.
Published: (2021)
by: Wu, Dongrui, et al.
Published: (2021)
SouLLMate: An Adaptive LLM-Driven System for Advanced Mental Health Support and Assessment, Based on a Systematic Application Survey
by: Guo, Qiming, et al.
Published: (2024)
by: Guo, Qiming, et al.
Published: (2024)
Measuring the Mental Health of Content Reviewers, a Systematic Review
by: Gonzalez, Alexandra, et al.
Published: (2025)
by: Gonzalez, Alexandra, et al.
Published: (2025)
Similar Items
-
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024) -
Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks
by: Guerdan, Luke, et al.
Published: (2025) -
Effects of Generative AI Errors on User Reliance Across Task Difficulty
by: Anthis, Jacy Reese, et al.
Published: (2026) -
Do Responsible AI Artifacts Advance Stakeholder Goals? Four Key Barriers Perceived by Legal and Civil Stakeholders
by: Kawakami, Anna, et al.
Published: (2024) -
Predictive Performance Comparison of Decision Policies Under Confounding
by: Guerdan, Luke, et al.
Published: (2024)