SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Dong, Li, Zhaowei, Wang, Pengyu, Zhang, Xin, Zhou, Yaqian, Qiu, Xipeng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917561426247680
author Zhang, Dong
Li, Zhaowei
Wang, Pengyu
Zhang, Xin
Zhou, Yaqian
Qiu, Xipeng
author_facet Zhang, Dong
Li, Zhaowei
Wang, Pengyu
Zhang, Xin
Zhou, Yaqian
Qiu, Xipeng
contents Human communication is a complex and diverse process that not only involves multiple factors such as language, commonsense, and cultural backgrounds but also requires the participation of multimodal information, such as speech. Large Language Model (LLM)-based multi-agent systems have demonstrated promising performance in simulating human society. Can we leverage LLM-based multi-agent systems to simulate human communication? However, current LLM-based multi-agent systems mainly rely on text as the primary medium. In this paper, we propose SpeechAgents, a multi-modal LLM based multi-agent system designed for simulating human communication. SpeechAgents utilizes multi-modal LLM as the control center for individual agent and employes multi-modal signals as the medium for exchanged messages among agents. Additionally, we propose Multi-Agent Tuning to enhance the multi-agent capabilities of LLM without compromising general abilities. To strengthen and evaluate the effectiveness of human communication simulation, we build the Human-Communication Simulation Benchmark. Experimental results demonstrate that SpeechAgents can simulate human communication dialogues with consistent content, authentic rhythm, and rich emotions and demonstrate excellent scalability even with up to 25 agents, which can apply to tasks such as drama creation and audio novels generation. Code and models will be open-sourced at https://github. com/0nutation/SpeechAgents
format Preprint
id arxiv_https___arxiv_org_abs_2401_03945
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems
Zhang, Dong
Li, Zhaowei
Wang, Pengyu
Zhang, Xin
Zhou, Yaqian
Qiu, Xipeng
Computation and Language
Human communication is a complex and diverse process that not only involves multiple factors such as language, commonsense, and cultural backgrounds but also requires the participation of multimodal information, such as speech. Large Language Model (LLM)-based multi-agent systems have demonstrated promising performance in simulating human society. Can we leverage LLM-based multi-agent systems to simulate human communication? However, current LLM-based multi-agent systems mainly rely on text as the primary medium. In this paper, we propose SpeechAgents, a multi-modal LLM based multi-agent system designed for simulating human communication. SpeechAgents utilizes multi-modal LLM as the control center for individual agent and employes multi-modal signals as the medium for exchanged messages among agents. Additionally, we propose Multi-Agent Tuning to enhance the multi-agent capabilities of LLM without compromising general abilities. To strengthen and evaluate the effectiveness of human communication simulation, we build the Human-Communication Simulation Benchmark. Experimental results demonstrate that SpeechAgents can simulate human communication dialogues with consistent content, authentic rhythm, and rich emotions and demonstrate excellent scalability even with up to 25 agents, which can apply to tasks such as drama creation and audio novels generation. Code and models will be open-sourced at https://github. com/0nutation/SpeechAgents
title SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems
topic Computation and Language
url https://arxiv.org/abs/2401.03945