Scaffolded Language Models with Language Supervision for Mixed-Autonomy: A Survey

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lin, Matthieu, Sheng, Jenny, Zhao, Andrew, Wang, Shenzhi, Yue, Yang, Huang, Victor Shea Jay, Liu, Huan, Liu, Jun, Huang, Gao, Liu, Yong-Jin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912687017951232
author Lin, Matthieu
Sheng, Jenny
Zhao, Andrew
Wang, Shenzhi
Yue, Yang
Huang, Victor Shea Jay
Liu, Huan
Liu, Jun
Huang, Gao
Liu, Yong-Jin
author_facet Lin, Matthieu
Sheng, Jenny
Zhao, Andrew
Wang, Shenzhi
Yue, Yang
Huang, Victor Shea Jay
Liu, Huan
Liu, Jun
Huang, Gao
Liu, Yong-Jin
contents This survey organizes the intricate literature on the design and optimization of emerging structures around post-trained LMs. We refer to this overarching structure as scaffolded LMs and focus on LMs that are integrated into multi-step processes with tools. We view scaffolded LMs as semi-parametric models wherein we train non-parametric variables, including the prompt, tools, and scaffold's code. In particular, they interpret instructions, use tools, and receive feedback all in language. Recent works use an LM as an optimizer to interpret language supervision and update non-parametric variables according to intricate objectives. In this survey, we refer to this paradigm as training of scaffolded LMs with language supervision. A key feature of non-parametric training is the ability to learn from language. Parametric training excels in learning from demonstration (supervised learning), exploration (reinforcement learning), or observations (unsupervised learning), using well-defined loss functions. Language-based optimization enables rich, interpretable, and expressive objectives, while mitigating issues like catastrophic forgetting and supporting compatibility with closed-source models. Furthermore, agents are increasingly deployed as co-workers in real-world applications such as Copilot in Office tools or software development. In these mixed-autonomy settings, where control and decision-making are shared between human and AI, users point out errors or suggest corrections. Accordingly, we discuss agents that continuously improve by learning from this real-time, language-based feedback and refer to this setting as streaming learning from language supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16392
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaffolded Language Models with Language Supervision for Mixed-Autonomy: A Survey
Lin, Matthieu
Sheng, Jenny
Zhao, Andrew
Wang, Shenzhi
Yue, Yang
Huang, Victor Shea Jay
Liu, Huan
Liu, Jun
Huang, Gao
Liu, Yong-Jin
Computation and Language
Machine Learning
This survey organizes the intricate literature on the design and optimization of emerging structures around post-trained LMs. We refer to this overarching structure as scaffolded LMs and focus on LMs that are integrated into multi-step processes with tools. We view scaffolded LMs as semi-parametric models wherein we train non-parametric variables, including the prompt, tools, and scaffold's code. In particular, they interpret instructions, use tools, and receive feedback all in language. Recent works use an LM as an optimizer to interpret language supervision and update non-parametric variables according to intricate objectives. In this survey, we refer to this paradigm as training of scaffolded LMs with language supervision. A key feature of non-parametric training is the ability to learn from language. Parametric training excels in learning from demonstration (supervised learning), exploration (reinforcement learning), or observations (unsupervised learning), using well-defined loss functions. Language-based optimization enables rich, interpretable, and expressive objectives, while mitigating issues like catastrophic forgetting and supporting compatibility with closed-source models. Furthermore, agents are increasingly deployed as co-workers in real-world applications such as Copilot in Office tools or software development. In these mixed-autonomy settings, where control and decision-making are shared between human and AI, users point out errors or suggest corrections. Accordingly, we discuss agents that continuously improve by learning from this real-time, language-based feedback and refer to this setting as streaming learning from language supervision.
title Scaffolded Language Models with Language Supervision for Mixed-Autonomy: A Survey
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.16392