EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ashktorab, Zahra, Geyer, Werner, Desmond, Michael, Daly, Elizabeth M., Cooper, Martin Santillan, Pan, Qian, Miehling, Erik, Pedapati, Tejaswini, Do, Hyo Jin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2024)
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2024)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
Human-Centered Design Recommendations for LLM-as-a-Judge
von: Pan, Qian, et al.
Veröffentlicht: (2024)
von: Pan, Qian, et al.
Veröffentlicht: (2024)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
von: Wagner, Nico, et al.
Veröffentlicht: (2024)
von: Wagner, Nico, et al.
Veröffentlicht: (2024)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
von: Chiang, Charles, et al.
Veröffentlicht: (2026)
von: Chiang, Charles, et al.
Veröffentlicht: (2026)
Granite Guardian
von: Padhi, Inkit, et al.
Veröffentlicht: (2024)
von: Padhi, Inkit, et al.
Veröffentlicht: (2024)
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2025)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2025)
Emerging Reliance Behaviors in Human-AI Content Grounded Data Generation: The Role of Cognitive Forcing Functions and Hallucinations
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2024)
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2024)
Interaction Configurations and Prompt Guidance in Conversational AI for Question Answering in Human-AI Teams
von: Song, Jaeyoon, et al.
Veröffentlicht: (2025)
von: Song, Jaeyoon, et al.
Veröffentlicht: (2025)
Who Sees the Risk? Stakeholder Conflicts and Explanatory Policies in LLM-based Risk Assessment
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
Hide or Highlight: Understanding the Impact of Factuality Expression on User Trust
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2026)
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2026)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
Localizing Persona Representations in LLMs
von: Cintas, Celia, et al.
Veröffentlicht: (2025)
von: Cintas, Celia, et al.
Veröffentlicht: (2025)
From PEFT to DEFT: Parameter Efficient Finetuning for Reducing Activation Density in Transformers
von: Runwal, Bharat, et al.
Veröffentlicht: (2024)
von: Runwal, Bharat, et al.
Veröffentlicht: (2024)
AI Steerability 360: A Toolkit for Steering Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2026)
von: Miehling, Erik, et al.
Veröffentlicht: (2026)
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
Facilitating Human-LLM Collaboration through Factuality Scores and Source Attributions
von: Do, Hyo Jin, et al.
Veröffentlicht: (2024)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2024)
Evaluating the Prompt Steerability of Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
LongFuncEval: Measuring the effectiveness of long context models for function calling
von: Kate, Kiran, et al.
Veröffentlicht: (2025)
von: Kate, Kiran, et al.
Veröffentlicht: (2025)
Helping the Helper: Supporting Peer Counselors via AI-Empowered Practice and Feedback
von: Hsu, Shang-Ling, et al.
Veröffentlicht: (2023)
von: Hsu, Shang-Ling, et al.
Veröffentlicht: (2023)
Modular Prompt Learning Improves Vision-Language Models
von: Huang, Zhenhan, et al.
Veröffentlicht: (2025)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2025)
Differentiable Prompt Learning for Vision Language Models
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
Intermediate Representations are Strong AI-Generated Image Detectors
von: Huang, Zhenhan, et al.
Veröffentlicht: (2026)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2026)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
von: Zhou, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zhou, Lingfeng, et al.
Veröffentlicht: (2025)
Large Language Model Confidence Estimation via Black-Box Access
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development
von: Geyer, Werner, et al.
Veröffentlicht: (2025)
von: Geyer, Werner, et al.
Veröffentlicht: (2025)
Graph is all you need? Lightweight data-agnostic neural architecture search without training
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
CollabEval: Enhancing LLM-as-a-Judge via Multi-Agent Collaboration
von: Qian, Yiyue, et al.
Veröffentlicht: (2026)
von: Qian, Yiyue, et al.
Veröffentlicht: (2026)
Togedule: Scheduling Meetings with Large Language Models and Adaptive Representations of Group Availability
von: Song, Jaeyoon, et al.
Veröffentlicht: (2025)
von: Song, Jaeyoon, et al.
Veröffentlicht: (2025)
CELL your Model: Contrastive Explanations for Large Language Models
von: Luss, Ronny, et al.
Veröffentlicht: (2024)
von: Luss, Ronny, et al.
Veröffentlicht: (2024)
Adaptive Trust Metrics for Multi-LLM Systems: Enhancing Reliability in Regulated Industries
von: Bollikonda, Tejaswini
Veröffentlicht: (2026)
von: Bollikonda, Tejaswini
Veröffentlicht: (2026)
Agentic AI Needs a Systems Theory
von: Miehling, Erik, et al.
Veröffentlicht: (2025)
von: Miehling, Erik, et al.
Veröffentlicht: (2025)
CoFrNets: Interpretable Neural Architecture Inspired by Continued Fractions
von: Puri, Isha, et al.
Veröffentlicht: (2025)
von: Puri, Isha, et al.
Veröffentlicht: (2025)
Sparse Gradient Compression for Fine-Tuning Large Language Models
von: Yang, David H., et al.
Veröffentlicht: (2025)
von: Yang, David H., et al.
Veröffentlicht: (2025)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
von: Pan, Tianjun, et al.
Veröffentlicht: (2026)
von: Pan, Tianjun, et al.
Veröffentlicht: (2026)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
Computer Assisted Projective Rigidity
von: Daly, Charles
Veröffentlicht: (2024)
von: Daly, Charles
Veröffentlicht: (2024)
Ähnliche Einträge
-
Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2024) -
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025) -
Human-Centered Design Recommendations for LLM-as-a-Judge
von: Pan, Qian, et al.
Veröffentlicht: (2024) -
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
von: Wagner, Nico, et al.
Veröffentlicht: (2024) -
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
von: Chiang, Charles, et al.
Veröffentlicht: (2026)