CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pandey, Punya Syon, Yang, Yongjin, Liu, Jiarui, Jin, Zhijing
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917287212089344
author Pandey, Punya Syon
Yang, Yongjin
Liu, Jiarui
Jin, Zhijing
author_facet Pandey, Punya Syon
Yang, Yongjin
Liu, Jiarui
Jin, Zhijing
contents Game-theoretic interactions between agents with Large Language Models (LLMs) have revealed many emergent capabilities, yet the linguistic diversity of these interactions has not been sufficiently quantified. In this paper, we present the Conversational Robustness Evaluation Score: CORE, a metric to quantify the effectiveness of language use within multi-agent systems across different game-theoretic interactions. CORE integrates measures of cluster entropy, lexical repetition, and semantic similarity, providing a direct lens of dialog quality. We apply CORE to pairwise LLM dialogs across competitive, cooperative, and neutral settings, further grounding our analysis in Zipf's and Heaps' Laws to characterize word frequency distributions and vocabulary growth. Our findings show that cooperative settings exhibit both steeper Zipf distributions and higher Heap exponents, indicating more repetition alongside greater vocabulary expansion. In contrast, competitive interactions display lower Zipf and Heaps exponents, reflecting less repetition and more constrained vocabularies. These results provide new insights into how social incentives influence language adaptation, and highlight CORE as a robust diagnostic for measuring linguistic robustness in multi-agent LLM systems. Our code is available at https://github.com/psyonp/core.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
Pandey, Punya Syon
Yang, Yongjin
Liu, Jiarui
Jin, Zhijing
Computation and Language
Artificial Intelligence
Machine Learning
Game-theoretic interactions between agents with Large Language Models (LLMs) have revealed many emergent capabilities, yet the linguistic diversity of these interactions has not been sufficiently quantified. In this paper, we present the Conversational Robustness Evaluation Score: CORE, a metric to quantify the effectiveness of language use within multi-agent systems across different game-theoretic interactions. CORE integrates measures of cluster entropy, lexical repetition, and semantic similarity, providing a direct lens of dialog quality. We apply CORE to pairwise LLM dialogs across competitive, cooperative, and neutral settings, further grounding our analysis in Zipf's and Heaps' Laws to characterize word frequency distributions and vocabulary growth. Our findings show that cooperative settings exhibit both steeper Zipf distributions and higher Heap exponents, indicating more repetition alongside greater vocabulary expansion. In contrast, competitive interactions display lower Zipf and Heaps exponents, reflecting less repetition and more constrained vocabularies. These results provide new insights into how social incentives influence language adaptation, and highlight CORE as a robust diagnostic for measuring linguistic robustness in multi-agent LLM systems. Our code is available at https://github.com/psyonp/core.
title CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.11915