Limits of Large Language Models in Debating Humans

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Flamino, James, Modi, Mohammed Shahid, Szymanski, Boleslaw K., Cross, Brendan, Mikolajczyk, Colton
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913673470017536
author Flamino, James
Modi, Mohammed Shahid
Szymanski, Boleslaw K.
Cross, Brendan
Mikolajczyk, Colton
author_facet Flamino, James
Modi, Mohammed Shahid
Szymanski, Boleslaw K.
Cross, Brendan
Mikolajczyk, Colton
contents Large Language Models (LLMs) have shown remarkable promise in communicating with humans. Their potential use as artificial partners with humans in sociological experiments involving conversation is an exciting prospect. But how viable is it? Here, we rigorously test the limits of agents that debate using LLMs in a preregistered study that runs multiple debate-based opinion consensus games. Each game starts with six humans, six agents, or three humans and three agents. We found that agents can blend in and concentrate on a debate's topic better than humans, improving the productivity of all players. Yet, humans perceive agents as less convincing and confident than other humans, and several behavioral metrics of humans and agents we collected deviate measurably from each other. We observed that agents are already decent debaters, but their behavior generates a pattern distinctly different from the human-generated data.
format Preprint
id arxiv_https___arxiv_org_abs_2402_06049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Limits of Large Language Models in Debating Humans
Flamino, James
Modi, Mohammed Shahid
Szymanski, Boleslaw K.
Cross, Brendan
Mikolajczyk, Colton
Artificial Intelligence
Computation and Language
Human-Computer Interaction
Applications
Large Language Models (LLMs) have shown remarkable promise in communicating with humans. Their potential use as artificial partners with humans in sociological experiments involving conversation is an exciting prospect. But how viable is it? Here, we rigorously test the limits of agents that debate using LLMs in a preregistered study that runs multiple debate-based opinion consensus games. Each game starts with six humans, six agents, or three humans and three agents. We found that agents can blend in and concentrate on a debate's topic better than humans, improving the productivity of all players. Yet, humans perceive agents as less convincing and confident than other humans, and several behavioral metrics of humans and agents we collected deviate measurably from each other. We observed that agents are already decent debaters, but their behavior generates a pattern distinctly different from the human-generated data.
title Limits of Large Language Models in Debating Humans
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
Applications
url https://arxiv.org/abs/2402.06049