Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sanayei, Reza, Vesic, Srdjan, Blanco, Eduardo, Surdeanu, Mihai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916957649895424
author Sanayei, Reza
Vesic, Srdjan
Blanco, Eduardo
Surdeanu, Mihai
author_facet Sanayei, Reza
Vesic, Srdjan
Blanco, Eduardo
Surdeanu, Mihai
contents Large Language Models (LLMs) excel at linear reasoning tasks but remain underexplored on non-linear structures such as those found in natural debates, which are best expressed as argument graphs. We evaluate whether LLMs can approximate structured reasoning from Computational Argumentation Theory (CAT). Specifically, we use Quantitative Argumentation Debate (QuAD) semantics, which assigns acceptability scores to arguments based on their attack and support relations. Given only dialogue-formatted debates from two NoDE datasets, models are prompted to rank arguments without access to the underlying graph. We test several LLMs under advanced instruction strategies, including Chain-of-Thought and In-Context Learning. While models show moderate alignment with QuAD rankings, performance degrades with longer inputs or disrupted discourse flow. Advanced prompting helps mitigate these effects by reducing biases related to argument length and position. Our findings highlight both the promise and limitations of LLMs in modeling formal argumentation semantics and motivate future work on graph-aware reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15739
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics
Sanayei, Reza
Vesic, Srdjan
Blanco, Eduardo
Surdeanu, Mihai
Computation and Language
Large Language Models (LLMs) excel at linear reasoning tasks but remain underexplored on non-linear structures such as those found in natural debates, which are best expressed as argument graphs. We evaluate whether LLMs can approximate structured reasoning from Computational Argumentation Theory (CAT). Specifically, we use Quantitative Argumentation Debate (QuAD) semantics, which assigns acceptability scores to arguments based on their attack and support relations. Given only dialogue-formatted debates from two NoDE datasets, models are prompted to rank arguments without access to the underlying graph. We test several LLMs under advanced instruction strategies, including Chain-of-Thought and In-Context Learning. While models show moderate alignment with QuAD rankings, performance degrades with longer inputs or disrupted discourse flow. Advanced prompting helps mitigate these effects by reducing biases related to argument length and position. Our findings highlight both the promise and limitations of LLMs in modeling formal argumentation semantics and motivate future work on graph-aware reasoning.
title Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics
topic Computation and Language
url https://arxiv.org/abs/2509.15739