The Comparative Trap: Pairwise Comparisons Amplifies Biased Preferences of LLM Evaluators
Fuente:
arXiv
Guardado en:
| Autores principales: | Jeong, Hawon, Park, ChaeHun, Hong, Jimin, Lee, Hojoon, Choo, Jaegul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison
por: Park, ChaeHun, et al.
Publicado: (2024)
por: Park, ChaeHun, et al.
Publicado: (2024)
Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
por: Park, ChaeHun, et al.
Publicado: (2024)
por: Park, ChaeHun, et al.
Publicado: (2024)
Breaking Chains: Unraveling the Links in Multi-Hop Knowledge Unlearning
por: Choi, Minseok, et al.
Publicado: (2024)
por: Choi, Minseok, et al.
Publicado: (2024)
Can Tool-augmented Large Language Models be Aware of Incomplete Conditions?
por: Yang, Seungbin, et al.
Publicado: (2024)
por: Yang, Seungbin, et al.
Publicado: (2024)
Not the Example, but the Process: How Self-Generated Examples Enhance LLM Reasoning
por: Gwak, Daehoon, et al.
Publicado: (2026)
por: Gwak, Daehoon, et al.
Publicado: (2026)
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration
por: Park, ChaeHun, et al.
Publicado: (2024)
por: Park, ChaeHun, et al.
Publicado: (2024)
Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs
por: Gwak, Daehoon, et al.
Publicado: (2025)
por: Gwak, Daehoon, et al.
Publicado: (2025)
LiveWeb-IE: A Benchmark For Online Web Information Extraction
por: Yang, Seungbin, et al.
Publicado: (2026)
por: Yang, Seungbin, et al.
Publicado: (2026)
Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering
por: Park, ChaeHun, et al.
Publicado: (2024)
por: Park, ChaeHun, et al.
Publicado: (2024)
Opt-Out: Investigating Entity-Level Unlearning for Large Language Models via Optimal Transport
por: Choi, Minseok, et al.
Publicado: (2024)
por: Choi, Minseok, et al.
Publicado: (2024)
Protecting Privacy Through Approximating Optimal Parameters for Sequence Unlearning in Language Models
por: Lee, Dohyun, et al.
Publicado: (2024)
por: Lee, Dohyun, et al.
Publicado: (2024)
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
por: Lee, Yunseung, et al.
Publicado: (2026)
por: Lee, Yunseung, et al.
Publicado: (2026)
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
por: Liusie, Adian, et al.
Publicado: (2024)
por: Liusie, Adian, et al.
Publicado: (2024)
Cross-Lingual Unlearning of Selective Knowledge in Multilingual Language Models
por: Choi, Minseok, et al.
Publicado: (2024)
por: Choi, Minseok, et al.
Publicado: (2024)
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
por: Liusie, Adian, et al.
Publicado: (2023)
por: Liusie, Adian, et al.
Publicado: (2023)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
por: Hwang, Dongyoon, et al.
Publicado: (2024)
por: Hwang, Dongyoon, et al.
Publicado: (2024)
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
por: Kim, Donghu, et al.
Publicado: (2024)
por: Kim, Donghu, et al.
Publicado: (2024)
Forecasting Future International Events: A Reliable Dataset for Text-Based Event Modeling
por: Gwak, Daehoon, et al.
Publicado: (2024)
por: Gwak, Daehoon, et al.
Publicado: (2024)
Comparing Developer and LLM Biases in Code Evaluation
por: Mittal, Aditya, et al.
Publicado: (2026)
por: Mittal, Aditya, et al.
Publicado: (2026)
Exploring In-context Example Generation for Machine Translation
por: Lee, Dohyun, et al.
Publicado: (2025)
por: Lee, Dohyun, et al.
Publicado: (2025)
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
por: Weissburg, Iain, et al.
Publicado: (2024)
por: Weissburg, Iain, et al.
Publicado: (2024)
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
por: Hwang, Dongyoon, et al.
Publicado: (2025)
por: Hwang, Dongyoon, et al.
Publicado: (2025)
Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
por: Park, Cheonbok, et al.
Publicado: (2025)
por: Park, Cheonbok, et al.
Publicado: (2025)
Building Resource-Constrained Language Agents: A Korean Case Study on Chemical Toxicity Information
por: Cho, Hojun, et al.
Publicado: (2025)
por: Cho, Hojun, et al.
Publicado: (2025)
From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review
por: Zhang, Yaohui, et al.
Publicado: (2025)
por: Zhang, Yaohui, et al.
Publicado: (2025)
PrefPO: Pairwise Preference Prompt Optimization
por: Singhal, Rahul, et al.
Publicado: (2026)
por: Singhal, Rahul, et al.
Publicado: (2026)
ExpGuard: LLM Content Moderation in Specialized Domains
por: Choi, Minseok, et al.
Publicado: (2026)
por: Choi, Minseok, et al.
Publicado: (2026)
Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages
por: Anand, Srija, et al.
Publicado: (2026)
por: Anand, Srija, et al.
Publicado: (2026)
Single Ground Truth Is Not Enough: Adding Flexibility to Aspect-Based Sentiment Analysis Evaluation
por: Yang, Soyoung, et al.
Publicado: (2024)
por: Yang, Soyoung, et al.
Publicado: (2024)
Retrieve Only Relevant Tables Whether Few or Many: Adaptive Table Retrieval Method
por: Kim, Taehee, et al.
Publicado: (2026)
por: Kim, Taehee, et al.
Publicado: (2026)
Style over Story: Measuring LLM Narrative Preferences via Structured Selection
por: Jung, Donghoon, et al.
Publicado: (2025)
por: Jung, Donghoon, et al.
Publicado: (2025)
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
por: Sandan, Isik Baran, et al.
Publicado: (2025)
por: Sandan, Isik Baran, et al.
Publicado: (2025)
Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
por: Jiang, Yuxin, et al.
Publicado: (2024)
por: Jiang, Yuxin, et al.
Publicado: (2024)
DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization
por: Deng, Mengyi, et al.
Publicado: (2026)
por: Deng, Mengyi, et al.
Publicado: (2026)
Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation
por: Jung, Kyudan, et al.
Publicado: (2025)
por: Jung, Kyudan, et al.
Publicado: (2025)
Models Know Models Best: Evaluation via Model-Preferred Formats
por: Lee, Joonhak, et al.
Publicado: (2026)
por: Lee, Joonhak, et al.
Publicado: (2026)
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
por: Liu, Yinhong, et al.
Publicado: (2024)
por: Liu, Yinhong, et al.
Publicado: (2024)
Evaluating the Consistency of LLM Evaluators
por: Lee, Noah, et al.
Publicado: (2024)
por: Lee, Noah, et al.
Publicado: (2024)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
por: Lawrence, Logan, et al.
Publicado: (2025)
por: Lawrence, Logan, et al.
Publicado: (2025)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
por: Lee, Dongryeol, et al.
Publicado: (2026)
por: Lee, Dongryeol, et al.
Publicado: (2026)
Ejemplares similares
-
PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison
por: Park, ChaeHun, et al.
Publicado: (2024) -
Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
por: Park, ChaeHun, et al.
Publicado: (2024) -
Breaking Chains: Unraveling the Links in Multi-Hop Knowledge Unlearning
por: Choi, Minseok, et al.
Publicado: (2024) -
Can Tool-augmented Large Language Models be Aware of Incomplete Conditions?
por: Yang, Seungbin, et al.
Publicado: (2024) -
Not the Example, but the Process: How Self-Generated Examples Enhance LLM Reasoning
por: Gwak, Daehoon, et al.
Publicado: (2026)