Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Sternlicht, Noy, Gera, Ariel, Bar-Haim, Roy, Hope, Tom, Slonim, Noam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation
by: Sternlicht, Noy, et al.
Published: (2025)
by: Sternlicht, Noy, et al.
Published: (2025)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Efficient Benchmarking of Language Models
by: Perlitz, Yotam, et al.
Published: (2023)
by: Perlitz, Yotam, et al.
Published: (2023)
Systematic Biases in LLM Simulations of Debates
by: Taubenfeld, Amir, et al.
Published: (2024)
by: Taubenfeld, Amir, et al.
Published: (2024)
Debatrix: Multi-dimensional Debate Judge with Iterative Chronological Analysis Based on LLM
by: Liang, Jingcong, et al.
Published: (2024)
by: Liang, Jingcong, et al.
Published: (2024)
Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics
by: Sanayei, Reza, et al.
Published: (2025)
by: Sanayei, Reza, et al.
Published: (2025)
In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis
by: Arnaout, Hiba, et al.
Published: (2025)
by: Arnaout, Hiba, et al.
Published: (2025)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)
by: Perlitz, Yotam, et al.
Published: (2024)
DebateQA: Evaluating Question Answering on Debatable Knowledge
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
Evaluating LLM-Driven Summarisation of Parliamentary Debates with Computational Argumentation
by: Cunningham, Eoghan, et al.
Published: (2026)
by: Cunningham, Eoghan, et al.
Published: (2026)
Debate Helps Weak Judges Reward Stronger Models
by: Elasky, Ethan, et al.
Published: (2026)
by: Elasky, Ethan, et al.
Published: (2026)
InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating
by: Wang, Fuyu, et al.
Published: (2025)
by: Wang, Fuyu, et al.
Published: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
by: Clegg, Kester, et al.
Published: (2025)
by: Clegg, Kester, et al.
Published: (2025)
A Debate-Driven Experiment on LLM Hallucinations and Accuracy
by: Li, Ray, et al.
Published: (2024)
by: Li, Ray, et al.
Published: (2024)
FinDebate: Multi-Agent Collaborative Intelligence for Financial Analysis
by: Cai, Tianshi, et al.
Published: (2025)
by: Cai, Tianshi, et al.
Published: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
by: Zhou, Yilun, et al.
Published: (2025)
by: Zhou, Yilun, et al.
Published: (2025)
iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM Inference
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
Latent Debate: A Surrogate Framework for Interpreting LLM Thinking
by: Chen, Lihu, et al.
Published: (2025)
by: Chen, Lihu, et al.
Published: (2025)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
by: Chen, Junjie, et al.
Published: (2026)
by: Chen, Junjie, et al.
Published: (2026)
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
by: Asad, Ali, et al.
Published: (2025)
by: Asad, Ali, et al.
Published: (2025)
Learning from Self-Debate: Preparing Reasoning Models for Multi-Agent Debate
by: Liu, Chenxi, et al.
Published: (2026)
by: Liu, Chenxi, et al.
Published: (2026)
Multiple LLM Agents Debate for Equitable Cultural Alignment
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Task-Adaptive Embedding Refinement via Test-time LLM Guidance
by: Gera, Ariel, et al.
Published: (2026)
by: Gera, Ariel, et al.
Published: (2026)
Debating Truth: Debate-driven Claim Verification with Multiple Large Language Model Agents
by: He, Haorui, et al.
Published: (2025)
by: He, Haorui, et al.
Published: (2025)
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
R-Debater: Retrieval-Augmented Debate Generation through Argumentative Memory
by: Li, Maoyuan, et al.
Published: (2025)
by: Li, Maoyuan, et al.
Published: (2025)
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
by: Yehudai, Asaf, et al.
Published: (2026)
by: Yehudai, Asaf, et al.
Published: (2026)
Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate
by: Zhang, Yiqun, et al.
Published: (2024)
by: Zhang, Yiqun, et al.
Published: (2024)
An Empirical Analysis on Large Language Models in Debate Evaluation
by: Liu, Xinyi, et al.
Published: (2024)
by: Liu, Xinyi, et al.
Published: (2024)
GroupDebate: Enhancing the Efficiency of Multi-Agent Debate Using Group Discussion
by: Liu, Tongxuan, et al.
Published: (2024)
by: Liu, Tongxuan, et al.
Published: (2024)
VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument Mining
by: Ruiz-Dolz, Ramon, et al.
Published: (2023)
by: Ruiz-Dolz, Ramon, et al.
Published: (2023)
Fool Me, Fool Me: User Attitudes Toward LLM Falsehoods
by: Nirman, Diana Bar-Or, et al.
Published: (2024)
by: Nirman, Diana Bar-Or, et al.
Published: (2024)
Social Reasoning in Machines: Investigating Collective Truth-Seeking Dynamics in Large Language Model Debate
by: Pecher, Tom
Published: (2026)
by: Pecher, Tom
Published: (2026)
Debate to Align: Reliable Entity Alignment through Two-Stage Multi-Agent Debate
by: Wang, Cunda, et al.
Published: (2026)
by: Wang, Cunda, et al.
Published: (2026)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution
by: Li, Han, et al.
Published: (2025)
by: Li, Han, et al.
Published: (2025)
Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
by: Chen, Xuhang, et al.
Published: (2025)
by: Chen, Xuhang, et al.
Published: (2025)
Similar Items
-
CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation
by: Sternlicht, Noy, et al.
Published: (2025) -
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024) -
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025) -
Efficient Benchmarking of Language Models
by: Perlitz, Yotam, et al.
Published: (2023) -
Systematic Biases in LLM Simulations of Debates
by: Taubenfeld, Amir, et al.
Published: (2024)