ArgBench: Benchmarking LLMs on Computational Argumentation Tasks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ajjour, Yamen, Quensel, Carlotta, Lipka, Nedim, Wachsmuth, Henning
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914488704303104
author Ajjour, Yamen
Quensel, Carlotta
Lipka, Nedim
Wachsmuth, Henning
author_facet Ajjour, Yamen
Quensel, Carlotta
Lipka, Nedim
Wachsmuth, Henning
contents Argumentation skills are an essential toolkit for large language models (LLMs). These skills are crucial in various use cases, including self-reflection, debating collaboratively for diverse answers, and countering hate speech. In this paper, we create the first benchmark for a standardized evaluation of LLM-based approaches to computational argumentation, encompassing 33 datasets from previous work in unified form. Using the benchmark, we evaluate the generalizability of five LLM families across 46 computational argumentation tasks that cover mining arguments, assessing perspectives, assessing argument quality, reasoning about arguments, and generating arguments. On the benchmark, we conduct an extensive systematic analysis of the contribution of few-shot examples, reasoning steps, model size, and training skills to the performance of LLMs on the computational argumentation tasks in the benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ArgBench: Benchmarking LLMs on Computational Argumentation Tasks
Ajjour, Yamen
Quensel, Carlotta
Lipka, Nedim
Wachsmuth, Henning
Computation and Language
Artificial Intelligence
Argumentation skills are an essential toolkit for large language models (LLMs). These skills are crucial in various use cases, including self-reflection, debating collaboratively for diverse answers, and countering hate speech. In this paper, we create the first benchmark for a standardized evaluation of LLM-based approaches to computational argumentation, encompassing 33 datasets from previous work in unified form. Using the benchmark, we evaluate the generalizability of five LLM families across 46 computational argumentation tasks that cover mining arguments, assessing perspectives, assessing argument quality, reasoning about arguments, and generating arguments. On the benchmark, we conduct an extensive systematic analysis of the contribution of few-shot examples, reasoning steps, model size, and training skills to the performance of LLMs on the computational argumentation tasks in the benchmark.
title ArgBench: Benchmarking LLMs on Computational Argumentation Tasks
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.17366