Language Model Networks: Supervision-Efficient Learning through Dense Communication

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Shiguang, Wang, Yaqing, Yao, Quanming
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917549799636992
author Wu, Shiguang
Wang, Yaqing
Yao, Quanming
author_facet Wu, Shiguang
Wang, Yaqing
Yao, Quanming
contents Language models are increasingly used not only as standalone predictors but also as components in larger inference systems, from test-time scaling to multi-agent collaboration. We study language model networks, where pre-trained language models serve as reusable nodes and intelligence emerges from their topology, communication, and optimization. Existing systems mostly communicate through natural language: easy to deploy, but discrete, inefficient, and hard to optimize from end-task supervision. We propose LMNet, a dense and differentiable realization of this paradigm. LMNet uses stripped LLMs as vertex modules and trainable seq2seq modules as communication edges, enabling intermediate nodes to exchange dense vectors while preserving natural-language input and output at the system boundary. By bypassing intermediate embedding and de-embedding, LMNet enables efficient information transfer, end-to-end gradient optimization, and learned communication beyond hand-designed protocols. Experiments show performance with small additional training cost and effective adaptation under limited supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12741
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Model Networks: Supervision-Efficient Learning through Dense Communication
Wu, Shiguang
Wang, Yaqing
Yao, Quanming
Artificial Intelligence
Language models are increasingly used not only as standalone predictors but also as components in larger inference systems, from test-time scaling to multi-agent collaboration. We study language model networks, where pre-trained language models serve as reusable nodes and intelligence emerges from their topology, communication, and optimization. Existing systems mostly communicate through natural language: easy to deploy, but discrete, inefficient, and hard to optimize from end-task supervision. We propose LMNet, a dense and differentiable realization of this paradigm. LMNet uses stripped LLMs as vertex modules and trainable seq2seq modules as communication edges, enabling intermediate nodes to exchange dense vectors while preserving natural-language input and output at the system boundary. By bypassing intermediate embedding and de-embedding, LMNet enables efficient information transfer, end-to-end gradient optimization, and learned communication beyond hand-designed protocols. Experiments show performance with small additional training cost and effective adaptation under limited supervision.
title Language Model Networks: Supervision-Efficient Learning through Dense Communication
topic Artificial Intelligence
url https://arxiv.org/abs/2505.12741