Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Shankar, Shreya, Zamfirescu-Pereira, J. D., Hartmann, Björn, Parameswaran, Aditya G., Arawjo, Ian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAG Without the Lag: Interactive Debugging for Retrieval-Augmented Generation Pipelines
by: Lauro, Quentin Romero, et al.
Published: (2025)
by: Lauro, Quentin Romero, et al.
Published: (2025)
ChainBuddy: An AI Agent System for Generating LLM Pipelines
by: Zhang, Jingyue, et al.
Published: (2024)
by: Zhang, Jingyue, et al.
Published: (2024)
Beyond Code Generation: LLM-supported Exploration of the Program Design Space
by: Zamfirescu-Pereira, J. D., et al.
Published: (2025)
by: Zamfirescu-Pereira, J. D., et al.
Published: (2025)
LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
by: Kuroki, So, et al.
Published: (2025)
by: Kuroki, So, et al.
Published: (2025)
ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
by: Arawjo, Ian, et al.
Published: (2023)
by: Arawjo, Ian, et al.
Published: (2023)
Who Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-Making
by: Chappidi, Shreya, et al.
Published: (2026)
by: Chappidi, Shreya, et al.
Published: (2026)
Low-Burden LLM-Based Preference Learning: Personalizing Assistive Robots from Natural Language Feedback for Users with Paralysis
by: Shankar, Keshav, et al.
Published: (2026)
by: Shankar, Keshav, et al.
Published: (2026)
Communication Styles and Reader Preferences of LLM and Human Experts in Explaining Health Information
by: Zhou, Jiawei, et al.
Published: (2025)
by: Zhou, Jiawei, et al.
Published: (2025)
Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making
by: Ma, Shuai, et al.
Published: (2024)
by: Ma, Shuai, et al.
Published: (2024)
Rambler: Supporting Writing With Speech via LLM-Assisted Gist Manipulation
by: Lin, Susan, et al.
Published: (2024)
by: Lin, Susan, et al.
Published: (2024)
"We Have No Idea How Models will Behave in Production until Production": How Engineers Operationalize Machine Learning
by: Shankar, Shreya, et al.
Published: (2024)
by: Shankar, Shreya, et al.
Published: (2024)
Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support
by: Pu, Kevin, et al.
Published: (2025)
by: Pu, Kevin, et al.
Published: (2025)
A Multi-Agent Conversational Bandit Approach to Online Evaluation and Selection of User-Aligned LLM Responses
by: Dai, Xiangxiang, et al.
Published: (2025)
by: Dai, Xiangxiang, et al.
Published: (2025)
LLM-Assisted Visual Analytics: Opportunities and Challenges
by: Hutchinson, Maeve, et al.
Published: (2024)
by: Hutchinson, Maeve, et al.
Published: (2024)
Challenges & Opportunities with LLM-Assisted Visualization Retargeting
by: Snyder, Luke S., et al.
Published: (2025)
by: Snyder, Luke S., et al.
Published: (2025)
Evaluating Human Trust in LLM-Based Planners: A Preliminary Study
by: Chen, Shenghui, et al.
Published: (2025)
by: Chen, Shenghui, et al.
Published: (2025)
FARPLS: A Feature-Augmented Robot Trajectory Preference Labeling System to Assist Human Labelers' Preference Elicitation
by: Lyu, Hanfang, et al.
Published: (2024)
by: Lyu, Hanfang, et al.
Published: (2024)
Fewer Than 1% of Explainable AI Papers Validate Explainability with Humans
by: Suh, Ashley, et al.
Published: (2025)
by: Suh, Ashley, et al.
Published: (2025)
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
by: Jung, Minji, et al.
Published: (2026)
by: Jung, Minji, et al.
Published: (2026)
Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making
by: Garcia, Basile, et al.
Published: (2024)
by: Garcia, Basile, et al.
Published: (2024)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
by: Do, Hyo Jin, et al.
Published: (2025)
by: Do, Hyo Jin, et al.
Published: (2025)
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments
by: Daynauth, Roland, et al.
Published: (2024)
by: Daynauth, Roland, et al.
Published: (2024)
Exploring LLM-Generated Feedback for Economics Essays: How Teaching Assistants Evaluate and Envision Its Use
by: Lu, Xinyi, et al.
Published: (2025)
by: Lu, Xinyi, et al.
Published: (2025)
Rambler in the Wild: A Diary Study of LLM-Assisted Writing With Speech
by: Yang, Xuyu, et al.
Published: (2025)
by: Yang, Xuyu, et al.
Published: (2025)
Steering Semantic Data Processing With DocWrangler
by: Shankar, Shreya, et al.
Published: (2025)
by: Shankar, Shreya, et al.
Published: (2025)
Assessing the Quality of Mental Health Support in LLM Responses through Multi-Attribute Human Evaluation
by: Badawi, Abeer, et al.
Published: (2026)
by: Badawi, Abeer, et al.
Published: (2026)
Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge
by: Elangovan, Aparna, et al.
Published: (2024)
by: Elangovan, Aparna, et al.
Published: (2024)
Pensieve Discuss: Scalable Small-Group CS Tutoring System with AI
by: Yang, Yoonseok, et al.
Published: (2024)
by: Yang, Yoonseok, et al.
Published: (2024)
Comparing Human Expertise and Large Language Models Embeddings in Content Validity Assessment of Personality Tests
by: Milano, Nicola, et al.
Published: (2025)
by: Milano, Nicola, et al.
Published: (2025)
Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
Dreaming to Assist: Learning to Align with Human Objectives for Shared Control in High-Speed Racing
by: DeCastro, Jonathan, et al.
Published: (2024)
by: DeCastro, Jonathan, et al.
Published: (2024)
Facilitating Human-LLM Collaboration through Factuality Scores and Source Attributions
by: Do, Hyo Jin, et al.
Published: (2024)
by: Do, Hyo Jin, et al.
Published: (2024)
Model Behavior Specification by Leveraging LLM Self-Playing and Self-Improving
by: Park, Soya, et al.
Published: (2025)
by: Park, Soya, et al.
Published: (2025)
Evaluating AI Alignment in LLMs: Output Analysis of Value Priorities Across 75 Models with Human Benchmarking
by: Lau, Gabriel Rongyang, et al.
Published: (2025)
by: Lau, Gabriel Rongyang, et al.
Published: (2025)
Human Decision-Making with Persuasive and Narrative LLM Explanations
by: Marusich, Laura R., et al.
Published: (2026)
by: Marusich, Laura R., et al.
Published: (2026)
Human-Centred LLM Privacy Audits: Findings and Frictions
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes
by: Chen, Junlong, et al.
Published: (2024)
by: Chen, Junlong, et al.
Published: (2024)
The Interaction Layer: An Exploration for Co-Designing User-LLM Interactions in Parental Wellbeing Support Systems
by: Viswanathan, Sruthi, et al.
Published: (2024)
by: Viswanathan, Sruthi, et al.
Published: (2024)
Measuring Successful Cooperation in Human-AI Teamwork: Development and Validation of the Perceived Cooperativity and Teaming Perception Scales
by: Attig, Christiane, et al.
Published: (2026)
by: Attig, Christiane, et al.
Published: (2026)
Similar Items
-
RAG Without the Lag: Interactive Debugging for Retrieval-Augmented Generation Pipelines
by: Lauro, Quentin Romero, et al.
Published: (2025) -
ChainBuddy: An AI Agent System for Generating LLM Pipelines
by: Zhang, Jingyue, et al.
Published: (2024) -
Beyond Code Generation: LLM-supported Exploration of the Program Design Space
by: Zamfirescu-Pereira, J. D., et al.
Published: (2025) -
LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
by: Kuroki, So, et al.
Published: (2025) -
ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
by: Arawjo, Ian, et al.
Published: (2023)