Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics
Fuente:
arXiv
Saved in:
| Main Authors: | Sanayei, Reza, Vesic, Srdjan, Blanco, Eduardo, Surdeanu, Mihai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation Extraction
by: Chakma, Aunabil, et al.
Published: (2026)
by: Chakma, Aunabil, et al.
Published: (2026)
Synthetic Dataset for Evaluating Complex Compositional Knowledge for Natural Language Inference
by: Akoju, Sushma Anand, et al.
Published: (2023)
by: Akoju, Sushma Anand, et al.
Published: (2023)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023)
by: Golchin, Shahriar, et al.
Published: (2023)
Automatic Debate Evaluation with Argumentation Semantics and Natural Language Argument Graph Networks
by: Ruiz-Dolz, Ramon, et al.
Published: (2022)
by: Ruiz-Dolz, Ramon, et al.
Published: (2022)
Can LLMs Extract Frame-Semantic Arguments?
by: Devasier, Jacob, et al.
Published: (2025)
by: Devasier, Jacob, et al.
Published: (2025)
Fane at SemEval-2025 Task 10: Zero-Shot Entity Framing with Large Language Models
by: Fane, Enfa, et al.
Published: (2025)
by: Fane, Enfa, et al.
Published: (2025)
Bridging the Long-Tail Gap: Robust Retrieval-Augmented Relation Completion via Multi-Stage Paraphrase Infusion
by: Alam, Fahmida, et al.
Published: (2026)
by: Alam, Fahmida, et al.
Published: (2026)
Impact Measures for Gradual Argumentation Semantics
by: Anaissy, Caren Al, et al.
Published: (2024)
by: Anaissy, Caren Al, et al.
Published: (2024)
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023)
by: Golchin, Shahriar, et al.
Published: (2023)
Rejecting Arguments Based on Doubt in Structured Bipolar Argumentation
by: Müller, Michael A., et al.
Published: (2026)
by: Müller, Michael A., et al.
Published: (2026)
Memorization in In-Context Learning
by: Golchin, Shahriar, et al.
Published: (2024)
by: Golchin, Shahriar, et al.
Published: (2024)
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026)
by: Islam, Md Asiful, et al.
Published: (2026)
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
by: Sternlicht, Noy, et al.
Published: (2025)
by: Sternlicht, Noy, et al.
Published: (2025)
Understanding Cultural Alignment in Multilingual LLMs via Natural Debate Statements
by: Negru, Vlad-Andrei, et al.
Published: (2026)
by: Negru, Vlad-Andrei, et al.
Published: (2026)
Towards Realistic Few-Shot Relation Extraction: A New Meta Dataset and Evaluation
by: Alam, Fahmida, et al.
Published: (2024)
by: Alam, Fahmida, et al.
Published: (2024)
Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation
by: Riaz, Haris, et al.
Published: (2025)
by: Riaz, Haris, et al.
Published: (2025)
Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization
by: Cooper, Portia, et al.
Published: (2024)
by: Cooper, Portia, et al.
Published: (2024)
Argument Collapse: LLMs Flatten Long-Form Public Debate
by: Kim, Yekyung, et al.
Published: (2026)
by: Kim, Yekyung, et al.
Published: (2026)
Evaluating LLM-Driven Summarisation of Parliamentary Debates with Computational Argumentation
by: Cunningham, Eoghan, et al.
Published: (2026)
by: Cunningham, Eoghan, et al.
Published: (2026)
ELLEN: Extremely Lightly Supervised Learning For Efficient Named Entity Recognition
by: Riaz, Haris, et al.
Published: (2024)
by: Riaz, Haris, et al.
Published: (2024)
On the Existence of an Inverse Solution for Preference-Based Reductions in Argumentation
by: Zaninotto, Alessio, et al.
Published: (2026)
by: Zaninotto, Alessio, et al.
Published: (2026)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
by: Thakur, Aman Singh, et al.
Published: (2024)
by: Thakur, Aman Singh, et al.
Published: (2024)
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
From Words to Numbers: Your Large Language Model Is Secretly A Capable Regressor When Given In-Context Examples
by: Vacareanu, Robert, et al.
Published: (2024)
by: Vacareanu, Robert, et al.
Published: (2024)
Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate
by: Zhang, Yiqun, et al.
Published: (2024)
by: Zhang, Yiqun, et al.
Published: (2024)
VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument Mining
by: Ruiz-Dolz, Ramon, et al.
Published: (2023)
by: Ruiz-Dolz, Ramon, et al.
Published: (2023)
R-Debater: Retrieval-Augmented Debate Generation through Argumentative Memory
by: Li, Maoyuan, et al.
Published: (2025)
by: Li, Maoyuan, et al.
Published: (2025)
LLMs as Debate Partners: Utilizing Genetic Algorithms and Adversarial Search for Adaptive Arguments
by: Aryan, Prakash
Published: (2024)
by: Aryan, Prakash
Published: (2024)
Peeking inside the Black-Box: Reinforcement Learning for Explainable and Accurate Relation Extraction
by: Guo, Xinyu, et al.
Published: (2025)
by: Guo, Xinyu, et al.
Published: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
by: Liu, Yixin, et al.
Published: (2026)
by: Liu, Yixin, et al.
Published: (2026)
Overview of AI-Debater 2023: The Challenges of Argument Generation Tasks
by: Lin, Jiayu, et al.
Published: (2024)
by: Lin, Jiayu, et al.
Published: (2024)
Enhancing Transformer RNNs with Multiple Temporal Perspectives
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
Asking and Answering Questions to Extract Event-Argument Structures
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
MorphNLI: A Stepwise Approach to Natural Language Inference Using Text Morphing
by: Negru, Vlad Andrei, et al.
Published: (2025)
by: Negru, Vlad Andrei, et al.
Published: (2025)
Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2025)
Best of Both Worlds: A Pliable and Generalizable Neuro-Symbolic Approach for Relation Classification
by: Vacareanu, Robert, et al.
Published: (2024)
by: Vacareanu, Robert, et al.
Published: (2024)
Can LLMs be Good Graph Judge for Knowledge Graph Construction?
by: Huang, Haoyu, et al.
Published: (2024)
by: Huang, Haoyu, et al.
Published: (2024)
Similar Items
-
Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation Extraction
by: Chakma, Aunabil, et al.
Published: (2026) -
Synthetic Dataset for Evaluating Complex Compositional Knowledge for Natural Language Inference
by: Akoju, Sushma Anand, et al.
Published: (2023) -
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
by: Golchin, Shahriar, et al.
Published: (2023) -
Automatic Debate Evaluation with Argumentation Semantics and Natural Language Argument Graph Networks
by: Ruiz-Dolz, Ramon, et al.
Published: (2022) -
Can LLMs Extract Frame-Semantic Arguments?
by: Devasier, Jacob, et al.
Published: (2025)