Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
Fuente:
arXiv
Saved in:
| Main Authors: | Do, Hyo Jin, Ashktorab, Zahra, Gajcin, Jasmina, Miehling, Erik, Cooper, Martín Santillán, Pan, Qian, Daly, Elizabeth M., Geyer, Werner |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
by: Ashktorab, Zahra, et al.
Published: (2025)
by: Ashktorab, Zahra, et al.
Published: (2025)
Human-Centered Design Recommendations for LLM-as-a-Judge
by: Pan, Qian, et al.
Published: (2024)
by: Pan, Qian, et al.
Published: (2024)
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
by: Ashktorab, Zahra, et al.
Published: (2024)
by: Ashktorab, Zahra, et al.
Published: (2024)
Who Sees the Risk? Stakeholder Conflicts and Explanatory Policies in LLM-based Risk Assessment
by: Yadav, Srishti, et al.
Published: (2025)
by: Yadav, Srishti, et al.
Published: (2025)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
by: Chiang, Charles, et al.
Published: (2026)
by: Chiang, Charles, et al.
Published: (2026)
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
by: Gajcin, Jasmina, et al.
Published: (2025)
by: Gajcin, Jasmina, et al.
Published: (2025)
Hide or Highlight: Understanding the Impact of Factuality Expression on User Trust
by: Do, Hyo Jin, et al.
Published: (2025)
by: Do, Hyo Jin, et al.
Published: (2025)
Interaction Configurations and Prompt Guidance in Conversational AI for Question Answering in Human-AI Teams
by: Song, Jaeyoon, et al.
Published: (2025)
by: Song, Jaeyoon, et al.
Published: (2025)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
by: Gebreegziabher, Simret Araya, et al.
Published: (2026)
by: Gebreegziabher, Simret Araya, et al.
Published: (2026)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
by: Wagner, Nico, et al.
Published: (2024)
by: Wagner, Nico, et al.
Published: (2024)
Emerging Reliance Behaviors in Human-AI Content Grounded Data Generation: The Role of Cognitive Forcing Functions and Hallucinations
by: Ashktorab, Zahra, et al.
Published: (2024)
by: Ashktorab, Zahra, et al.
Published: (2024)
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
by: Do, Hyo Jin, et al.
Published: (2025)
by: Do, Hyo Jin, et al.
Published: (2025)
Facilitating Human-LLM Collaboration through Factuality Scores and Source Attributions
by: Do, Hyo Jin, et al.
Published: (2024)
by: Do, Hyo Jin, et al.
Published: (2024)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Helping the Helper: Supporting Peer Counselors via AI-Empowered Practice and Feedback
by: Hsu, Shang-Ling, et al.
Published: (2023)
by: Hsu, Shang-Ling, et al.
Published: (2023)
Togedule: Scheduling Meetings with Large Language Models and Adaptive Representations of Group Availability
by: Song, Jaeyoon, et al.
Published: (2025)
by: Song, Jaeyoon, et al.
Published: (2025)
A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development
by: Geyer, Werner, et al.
Published: (2025)
by: Geyer, Werner, et al.
Published: (2025)
Evaluating the Prompt Steerability of Large Language Models
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
DesignerlyLoop: Forming Design Intent through Curated Reasoning for Human-LLM Alignment
by: Wang, Anqi, et al.
Published: (2025)
by: Wang, Anqi, et al.
Published: (2025)
Human-in-the-Loop Synthetic Text Data Inspection with Provenance Tracking
by: Kang, Hong Jin, et al.
Published: (2024)
by: Kang, Hong Jin, et al.
Published: (2024)
Explainable Iterative Data Visualisation Refinement via an LLM Agent
by: Susam, Burak, et al.
Published: (2026)
by: Susam, Burak, et al.
Published: (2026)
Redefining Counterfactual Explanations for Reinforcement Learning: Overview, Challenges and Opportunities
by: Gajcin, Jasmina, et al.
Published: (2022)
by: Gajcin, Jasmina, et al.
Published: (2022)
ACTER: Diverse and Actionable Counterfactual Sequences for Explaining and Diagnosing RL Policies
by: Gajcin, Jasmina, et al.
Published: (2024)
by: Gajcin, Jasmina, et al.
Published: (2024)
Who's Sorry Now: User Preferences Among Rote, Empathic, and Explanatory Apologies from LLM Chatbots
by: Ashktorab, Zahra, et al.
Published: (2025)
by: Ashktorab, Zahra, et al.
Published: (2025)
SLInterpreter: An Exploratory and Iterative Human-AI Collaborative System for GNN-based Synthetic Lethal Prediction
by: Jiang, Haoran, et al.
Published: (2024)
by: Jiang, Haoran, et al.
Published: (2024)
Agentic AI Needs a Systems Theory
by: Miehling, Erik, et al.
Published: (2025)
by: Miehling, Erik, et al.
Published: (2025)
DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection
by: Tang, Yuying, et al.
Published: (2026)
by: Tang, Yuying, et al.
Published: (2026)
Current and Future Use of Large Language Models for Knowledge Work
by: Brachman, Michelle, et al.
Published: (2025)
by: Brachman, Michelle, et al.
Published: (2025)
Exploring the Impact of an LLM-Powered Teachable Agent on Learning Gains and Cognitive Load in Music Education
by: Jin, Lingxi, et al.
Published: (2025)
by: Jin, Lingxi, et al.
Published: (2025)
Beyond Visualization: Building Decision Intelligence Through Iterative Dashboard Refinement
by: Tadakala, Likitha, et al.
Published: (2025)
by: Tadakala, Likitha, et al.
Published: (2025)
ColorCode: A Bayesian Approach to Augmentative and Alternative Communication with Two Buttons
by: Daly, Matthew
Published: (2022)
by: Daly, Matthew
Published: (2022)
"The Diagram is like Guardrails": Structuring GenAI-assisted Hypotheses Exploration with an Interactive Shared Representation
by: Ding, Zijian, et al.
Published: (2025)
by: Ding, Zijian, et al.
Published: (2025)
Continual Human-in-the-Loop Optimization
by: Liao, Yi-Chi, et al.
Published: (2025)
by: Liao, Yi-Chi, et al.
Published: (2025)
"Better Ask for Forgiveness than Permission": Practices and Policies of AI Disclosure in Freelance Work
by: Hwang, Angel Hsing-Chi, et al.
Published: (2026)
by: Hwang, Angel Hsing-Chi, et al.
Published: (2026)
Schemex: Discovering Design Patterns from Examples through Iterative Abstraction and Refinement
by: Wang, Sitong, et al.
Published: (2025)
by: Wang, Sitong, et al.
Published: (2025)
Environment-Aware and Human-Cooperative Swing Control for Lower-Limb Prostheses in Diverse Obstacle Scenarios
by: Xing, Haosen, et al.
Published: (2025)
by: Xing, Haosen, et al.
Published: (2025)
Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge
by: Elangovan, Aparna, et al.
Published: (2024)
by: Elangovan, Aparna, et al.
Published: (2024)
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
by: Szymanski, Annalisa, et al.
Published: (2024)
by: Szymanski, Annalisa, et al.
Published: (2024)
CogInstrument: Modeling Cognitive Processes for Bidirectional Human-LLM Alignment in Planning Tasks
by: Wang, Anqi, et al.
Published: (2026)
by: Wang, Anqi, et al.
Published: (2026)
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
by: Nogueira, Brenda, et al.
Published: (2025)
by: Nogueira, Brenda, et al.
Published: (2025)
Similar Items
-
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
by: Ashktorab, Zahra, et al.
Published: (2025) -
Human-Centered Design Recommendations for LLM-as-a-Judge
by: Pan, Qian, et al.
Published: (2024) -
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
by: Ashktorab, Zahra, et al.
Published: (2024) -
Who Sees the Risk? Stakeholder Conflicts and Explanatory Policies in LLM-based Risk Assessment
by: Yadav, Srishti, et al.
Published: (2025) -
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
by: Chiang, Charles, et al.
Published: (2026)