Humans or LLMs as the Judge? A Study on Judgement Biases
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Guiming Hardy, Chen, Shunian, Liu, Ziche, Jiang, Feng, Wang, Benyou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MileBench: Benchmarking MLLMs in Long Context
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024)
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024)
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
von: Ge, Wentao, et al.
Veröffentlicht: (2023)
von: Ge, Wentao, et al.
Veröffentlicht: (2023)
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
CMB: A Comprehensive Medical Benchmark in Chinese
von: Wang, Xidong, et al.
Veröffentlicht: (2023)
von: Wang, Xidong, et al.
Veröffentlicht: (2023)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language Models
von: Liu, Ziche, et al.
Veröffentlicht: (2024)
von: Liu, Ziche, et al.
Veröffentlicht: (2024)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
von: Chen, Junying, et al.
Veröffentlicht: (2023)
von: Chen, Junying, et al.
Veröffentlicht: (2023)
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
LLM4PM: A case study on using Large Language Models for Process Modeling in Enterprise Organizations
von: Ziche, Clara, et al.
Veröffentlicht: (2024)
von: Ziche, Clara, et al.
Veröffentlicht: (2024)
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2024)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2024)
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
von: Zhang, Bingquan, et al.
Veröffentlicht: (2025)
von: Zhang, Bingquan, et al.
Veröffentlicht: (2025)
LLMs for Doctors: Leveraging Medical LLMs to Assist Doctors, Not Replace Them
von: Xie, Wenya, et al.
Veröffentlicht: (2024)
von: Xie, Wenya, et al.
Veröffentlicht: (2024)
PlatoLM: Teaching LLMs in Multi-Round Dialogue via a User Simulator
von: Kong, Chuyi, et al.
Veröffentlicht: (2023)
von: Kong, Chuyi, et al.
Veröffentlicht: (2023)
Evaluating Gender Bias of LLMs in Making Morality Judgements
von: Bajaj, Divij, et al.
Veröffentlicht: (2024)
von: Bajaj, Divij, et al.
Veröffentlicht: (2024)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
Do LLMs Triage Like Clinicians? A Dynamic Study of Outpatient Referral
von: Liu, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxiao, et al.
Veröffentlicht: (2025)
LLMs Could Autonomously Learn Without External Supervision
von: Ji, Ke, et al.
Veröffentlicht: (2024)
von: Ji, Ke, et al.
Veröffentlicht: (2024)
Curse of Knowledge: When Complex Evaluation Context Benefits yet Biases LLM Judges
von: Li, Weiyuan, et al.
Veröffentlicht: (2025)
von: Li, Weiyuan, et al.
Veröffentlicht: (2025)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
von: Liu, Wanlong, et al.
Veröffentlicht: (2024)
von: Liu, Wanlong, et al.
Veröffentlicht: (2024)
Smurfs: Multi-Agent System using Context-Efficient DFSDT for Tool Planning
von: Chen, Junzhi, et al.
Veröffentlicht: (2024)
von: Chen, Junzhi, et al.
Veröffentlicht: (2024)
Direct Judgement Preference Optimization
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts
von: Zheng, Guorui, et al.
Veröffentlicht: (2024)
von: Zheng, Guorui, et al.
Veröffentlicht: (2024)
Measuring Teaching with LLMs
von: Hardy, Michael
Veröffentlicht: (2025)
von: Hardy, Michael
Veröffentlicht: (2025)
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
von: Martynov, Nikita, et al.
Veröffentlicht: (2025)
von: Martynov, Nikita, et al.
Veröffentlicht: (2025)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
von: Yang, Gao, et al.
Veröffentlicht: (2025)
von: Yang, Gao, et al.
Veröffentlicht: (2025)
Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
von: Xie, Wenya, et al.
Veröffentlicht: (2025)
von: Xie, Wenya, et al.
Veröffentlicht: (2025)
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
von: Bavaresco, Anna, et al.
Veröffentlicht: (2024)
von: Bavaresco, Anna, et al.
Veröffentlicht: (2024)
Incorporating Precedents for Legal Judgement Prediction on European Court of Human Rights Cases
von: Santosh, T. Y. S. S., et al.
Veröffentlicht: (2024)
von: Santosh, T. Y. S. S., et al.
Veröffentlicht: (2024)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement Measurement
von: Cheng, Zihao, et al.
Veröffentlicht: (2024)
von: Cheng, Zihao, et al.
Veröffentlicht: (2024)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
von: Lee, Sua, et al.
Veröffentlicht: (2026)
von: Lee, Sua, et al.
Veröffentlicht: (2026)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
von: Moon, Jiwon, et al.
Veröffentlicht: (2025)
von: Moon, Jiwon, et al.
Veröffentlicht: (2025)
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
von: Huang, Hui, et al.
Veröffentlicht: (2026)
von: Huang, Hui, et al.
Veröffentlicht: (2026)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MileBench: Benchmarking MLLMs in Long Context
von: Song, Dingjie, et al.
Veröffentlicht: (2024) -
ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
von: Chen, Guiming Hardy, et al.
Veröffentlicht: (2024) -
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
von: Chen, Junying, et al.
Veröffentlicht: (2024) -
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
von: Ge, Wentao, et al.
Veröffentlicht: (2023) -
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
von: Song, Dingjie, et al.
Veröffentlicht: (2024)