Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Churina, Svetlana, Jaidka, Kokil
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914128629596160
author Churina, Svetlana
Jaidka, Kokil
author_facet Churina, Svetlana
Jaidka, Kokil
contents Incivility on platforms such as Twitter (now X) and Reddit complicates the development of AI systems that can support productive, rhetorically sound political argumentation. We present experiments with \textit{GPT-3.5 Turbo} fine-tuned on two contrasting datasets of political discourse: high-incivility Twitter replies to U.S. Congress and low-incivility posts from Reddit's \textit{r/ChangeMyView}. Our evaluation examines how data composition and prompting strategies affect the rhetorical framing and deliberative quality of model-generated arguments. Results show that Reddit-finetuned models generate safer but rhetorically rigid arguments, while cross-platform fine-tuning amplifies adversarial tone and toxicity. Prompt-based steering reduces overt toxicity (e.g., personal attacks) but cannot fully offset the influence of noisy training data. We introduce a rhetorical evaluation rubric - covering justification, reciprocity, alignment, and authority - and provide implementation guidelines for authoring, moderation, and deliberation-support systems.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16813
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation
Churina, Svetlana
Jaidka, Kokil
Computation and Language
Artificial Intelligence
Incivility on platforms such as Twitter (now X) and Reddit complicates the development of AI systems that can support productive, rhetorically sound political argumentation. We present experiments with \textit{GPT-3.5 Turbo} fine-tuned on two contrasting datasets of political discourse: high-incivility Twitter replies to U.S. Congress and low-incivility posts from Reddit's \textit{r/ChangeMyView}. Our evaluation examines how data composition and prompting strategies affect the rhetorical framing and deliberative quality of model-generated arguments. Results show that Reddit-finetuned models generate safer but rhetorically rigid arguments, while cross-platform fine-tuning amplifies adversarial tone and toxicity. Prompt-based steering reduces overt toxicity (e.g., personal attacks) but cannot fully offset the influence of noisy training data. We introduce a rhetorical evaluation rubric - covering justification, reciprocity, alignment, and authority - and provide implementation guidelines for authoring, moderation, and deliberation-support systems.
title Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.16813