GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lee, Jisoo, Chang, Raeyoung, Kwon, Dongwook, Singh, Harmanpreet, Verma, Nikhil
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913946713194496
author Lee, Jisoo
Chang, Raeyoung
Kwon, Dongwook
Singh, Harmanpreet
Verma, Nikhil
author_facet Lee, Jisoo
Chang, Raeyoung
Kwon, Dongwook
Singh, Harmanpreet
Verma, Nikhil
contents Multi-agent systems built on language models have shown strong performance on collaborative reasoning tasks. However, existing evaluations focus only on the correctness of the final output, overlooking how inefficient communication and poor coordination contribute to redundant reasoning and higher computational costs. We introduce GEMMAS, a graph-based evaluation framework that analyzes the internal collaboration process by modeling agent interactions as a directed acyclic graph. To capture collaboration quality, we propose two process-level metrics: Information Diversity Score (IDS) to measure semantic variation in inter-agent messages, and Unnecessary Path Ratio (UPR) to quantify redundant reasoning paths. We evaluate GEMMAS across five benchmarks and highlight results on GSM8K, where systems with only a 2.1% difference in accuracy differ by 12.8% in IDS and 80% in UPR, revealing substantial variation in internal collaboration. These findings demonstrate that outcome-only metrics are insufficient for evaluating multi-agent performance and highlight the importance of process-level diagnostics in designing more interpretable and resource-efficient collaborative AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13190
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems
Lee, Jisoo
Chang, Raeyoung
Kwon, Dongwook
Singh, Harmanpreet
Verma, Nikhil
Computation and Language
Multi-agent systems built on language models have shown strong performance on collaborative reasoning tasks. However, existing evaluations focus only on the correctness of the final output, overlooking how inefficient communication and poor coordination contribute to redundant reasoning and higher computational costs. We introduce GEMMAS, a graph-based evaluation framework that analyzes the internal collaboration process by modeling agent interactions as a directed acyclic graph. To capture collaboration quality, we propose two process-level metrics: Information Diversity Score (IDS) to measure semantic variation in inter-agent messages, and Unnecessary Path Ratio (UPR) to quantify redundant reasoning paths. We evaluate GEMMAS across five benchmarks and highlight results on GSM8K, where systems with only a 2.1% difference in accuracy differ by 12.8% in IDS and 80% in UPR, revealing substantial variation in internal collaboration. These findings demonstrate that outcome-only metrics are insufficient for evaluating multi-agent performance and highlight the importance of process-level diagnostics in designing more interpretable and resource-efficient collaborative AI systems.
title GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems
topic Computation and Language
url https://arxiv.org/abs/2507.13190