Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ashktorab, Zahra, Desmond, Michael, Pan, Qian, Johnson, James M., Cooper, Martin Santillan, Daly, Elizabeth M., Nair, Rahul, Pedapati, Tejaswini, Do, Hyo Jin, Geyer, Werner |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2025)
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2025)
Human-Centered Design Recommendations for LLM-as-a-Judge
von: Pan, Qian, et al.
Veröffentlicht: (2024)
von: Pan, Qian, et al.
Veröffentlicht: (2024)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
von: Wagner, Nico, et al.
Veröffentlicht: (2024)
von: Wagner, Nico, et al.
Veröffentlicht: (2024)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
von: Chiang, Charles, et al.
Veröffentlicht: (2026)
von: Chiang, Charles, et al.
Veröffentlicht: (2026)
Hide or Highlight: Understanding the Impact of Factuality Expression on User Trust
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
Emerging Reliance Behaviors in Human-AI Content Grounded Data Generation: The Role of Cognitive Forcing Functions and Hallucinations
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2024)
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2024)
Granite Guardian
von: Padhi, Inkit, et al.
Veröffentlicht: (2024)
von: Padhi, Inkit, et al.
Veröffentlicht: (2024)
Interaction Configurations and Prompt Guidance in Conversational AI for Question Answering in Human-AI Teams
von: Song, Jaeyoon, et al.
Veröffentlicht: (2025)
von: Song, Jaeyoon, et al.
Veröffentlicht: (2025)
From PEFT to DEFT: Parameter Efficient Finetuning for Reducing Activation Density in Transformers
von: Runwal, Bharat, et al.
Veröffentlicht: (2024)
von: Runwal, Bharat, et al.
Veröffentlicht: (2024)
LongFuncEval: Measuring the effectiveness of long context models for function calling
von: Kate, Kiran, et al.
Veröffentlicht: (2025)
von: Kate, Kiran, et al.
Veröffentlicht: (2025)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2026)
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2026)
Helping the Helper: Supporting Peer Counselors via AI-Empowered Practice and Feedback
von: Hsu, Shang-Ling, et al.
Veröffentlicht: (2023)
von: Hsu, Shang-Ling, et al.
Veröffentlicht: (2023)
Modular Prompt Learning Improves Vision-Language Models
von: Huang, Zhenhan, et al.
Veröffentlicht: (2025)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2025)
Differentiable Prompt Learning for Vision Language Models
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
Intermediate Representations are Strong AI-Generated Image Detectors
von: Huang, Zhenhan, et al.
Veröffentlicht: (2026)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2026)
Large Language Model Confidence Estimation via Black-Box Access
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development
von: Geyer, Werner, et al.
Veröffentlicht: (2025)
von: Geyer, Werner, et al.
Veröffentlicht: (2025)
Graph is all you need? Lightweight data-agnostic neural architecture search without training
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2024)
Aligning Judgment Using Task Context and Explanations to Improve Human-Recommender System Performance
von: Srivastava, Divya, et al.
Veröffentlicht: (2024)
von: Srivastava, Divya, et al.
Veröffentlicht: (2024)
Togedule: Scheduling Meetings with Large Language Models and Adaptive Representations of Group Availability
von: Song, Jaeyoon, et al.
Veröffentlicht: (2025)
von: Song, Jaeyoon, et al.
Veröffentlicht: (2025)
ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical Judgment
von: Li, Ruochen, et al.
Veröffentlicht: (2025)
von: Li, Ruochen, et al.
Veröffentlicht: (2025)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
von: Zhou, Han, et al.
Veröffentlicht: (2024)
von: Zhou, Han, et al.
Veröffentlicht: (2024)
CoFrNets: Interpretable Neural Architecture Inspired by Continued Fractions
von: Puri, Isha, et al.
Veröffentlicht: (2025)
von: Puri, Isha, et al.
Veröffentlicht: (2025)
Sparse Gradient Compression for Fine-Tuning Large Language Models
von: Yang, David H., et al.
Veröffentlicht: (2025)
von: Yang, David H., et al.
Veröffentlicht: (2025)
"The Diagram is like Guardrails": Structuring GenAI-assisted Hypotheses Exploration with an Interactive Shared Representation
von: Ding, Zijian, et al.
Veröffentlicht: (2025)
von: Ding, Zijian, et al.
Veröffentlicht: (2025)
A Representation-Level Assessment of Bias Mitigation in Foundation Models
von: Nizhnichenkov, Svetoslav, et al.
Veröffentlicht: (2026)
von: Nizhnichenkov, Svetoslav, et al.
Veröffentlicht: (2026)
Computer Assisted Projective Rigidity
von: Daly, Charles
Veröffentlicht: (2024)
von: Daly, Charles
Veröffentlicht: (2024)
TOPJoin: A Context-Aware Multi-Criteria Approach for Joinable Column Search
von: Kokel, Harsha, et al.
Veröffentlicht: (2025)
von: Kokel, Harsha, et al.
Veröffentlicht: (2025)
Evaluating Joinable Column Discovery Approaches for Context-Aware Search
von: Kokel, Harsha, et al.
Veröffentlicht: (2025)
von: Kokel, Harsha, et al.
Veröffentlicht: (2025)
NeuroPrune: A Neuro-inspired Topological Sparse Training Algorithm for Large Language Models
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
Paying Alignment Tax with Contrastive Learning
von: Korkmaz, Buse Sibel, et al.
Veröffentlicht: (2025)
von: Korkmaz, Buse Sibel, et al.
Veröffentlicht: (2025)
Facilitating Human-LLM Collaboration through Factuality Scores and Source Attributions
von: Do, Hyo Jin, et al.
Veröffentlicht: (2024)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2024)
Humble AI in the real-world: the case of algorithmic hiring
von: Nair, Rahul, et al.
Veröffentlicht: (2025)
von: Nair, Rahul, et al.
Veröffentlicht: (2025)
Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change
von: Kurfalı, Murathan, et al.
Veröffentlicht: (2025)
von: Kurfalı, Murathan, et al.
Veröffentlicht: (2025)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Aligning ASR Evaluation with Human and LLM Judgments: Intelligibility Metrics Using Phonetic, Semantic, and NLI Approaches
von: Phukon, Bornali, et al.
Veröffentlicht: (2025)
von: Phukon, Bornali, et al.
Veröffentlicht: (2025)
Reasons to Reject? Aligning Language Models with Judgments
von: Xu, Weiwen, et al.
Veröffentlicht: (2023)
von: Xu, Weiwen, et al.
Veröffentlicht: (2023)
Ranking Large Language Models without Ground Truth
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2025) -
Human-Centered Design Recommendations for LLM-as-a-Judge
von: Pan, Qian, et al.
Veröffentlicht: (2024) -
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
von: Wagner, Nico, et al.
Veröffentlicht: (2024) -
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025) -
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
von: Chiang, Charles, et al.
Veröffentlicht: (2026)