Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Jae Hee, Lauscher, Anne, Albrecht, Stefano V.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914179767599104
author Lee, Jae Hee
Lauscher, Anne
Albrecht, Stefano V.
author_facet Lee, Jae Hee
Lauscher, Anne
Albrecht, Stefano V.
contents Large language models (LLMs) have been widely deployed in various applications, often functioning as autonomous agents that interact with each other in multi-agent systems. While these systems have shown promise in enhancing capabilities and enabling complex tasks, they also pose significant ethical challenges. This position paper outlines a research agenda aimed at ensuring the ethical behavior of multi-agent systems of LLMs (MALMs) from the perspective of mechanistic interpretability. We identify three key research challenges: (i) developing comprehensive evaluation frameworks to assess ethical behavior at individual, interactional, and systemic levels; (ii) elucidating the internal mechanisms that give rise to emergent behaviors through mechanistic interpretability; and (iii) implementing targeted parameter-efficient alignment techniques to steer MALMs towards ethical behaviors without compromising their performance.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04691
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
Lee, Jae Hee
Lauscher, Anne
Albrecht, Stefano V.
Artificial Intelligence
Computation and Language
Multiagent Systems
Large language models (LLMs) have been widely deployed in various applications, often functioning as autonomous agents that interact with each other in multi-agent systems. While these systems have shown promise in enhancing capabilities and enabling complex tasks, they also pose significant ethical challenges. This position paper outlines a research agenda aimed at ensuring the ethical behavior of multi-agent systems of LLMs (MALMs) from the perspective of mechanistic interpretability. We identify three key research challenges: (i) developing comprehensive evaluation frameworks to assess ethical behavior at individual, interactional, and systemic levels; (ii) elucidating the internal mechanisms that give rise to emergent behaviors through mechanistic interpretability; and (iii) implementing targeted parameter-efficient alignment techniques to steer MALMs towards ethical behaviors without compromising their performance.
title Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
topic Artificial Intelligence
Computation and Language
Multiagent Systems
url https://arxiv.org/abs/2512.04691