Retrospective Learning from Interactions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zizhao, Gul, Mustafa Omer, Chen, Yiwei, Geng, Gloria, Wu, Anne, Artzi, Yoav
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910958182465536
author Chen, Zizhao
Gul, Mustafa Omer
Chen, Yiwei
Geng, Gloria
Wu, Anne
Artzi, Yoav
author_facet Chen, Zizhao
Gul, Mustafa Omer
Chen, Yiwei
Geng, Gloria
Wu, Anne
Artzi, Yoav
contents Multi-turn interactions between large language models (LLMs) and users naturally include implicit feedback signals. If an LLM responds in an unexpected way to an instruction, the user is likely to signal it by rephrasing the request, expressing frustration, or pivoting to an alternative task. Such signals are task-independent and occupy a relatively constrained subspace of language, allowing the LLM to identify them even if it fails on the actual task. We introduce ReSpect, a method to learn from such signals in past interactions via retrospection without additional annotations. We deploy ReSpect in a new multimodal interaction scenario, where humans instruct a multimodal LLM to solve an abstract reasoning task with a combinatorial solution space. Through thousands of interactions with humans, we show how ReSpect gradually improves task completion rate from 31% to 82%, all without any external annotation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13852
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Retrospective Learning from Interactions
Chen, Zizhao
Gul, Mustafa Omer
Chen, Yiwei
Geng, Gloria
Wu, Anne
Artzi, Yoav
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Multi-turn interactions between large language models (LLMs) and users naturally include implicit feedback signals. If an LLM responds in an unexpected way to an instruction, the user is likely to signal it by rephrasing the request, expressing frustration, or pivoting to an alternative task. Such signals are task-independent and occupy a relatively constrained subspace of language, allowing the LLM to identify them even if it fails on the actual task. We introduce ReSpect, a method to learn from such signals in past interactions via retrospection without additional annotations. We deploy ReSpect in a new multimodal interaction scenario, where humans instruct a multimodal LLM to solve an abstract reasoning task with a combinatorial solution space. Through thousands of interactions with humans, we show how ReSpect gradually improves task completion rate from 31% to 82%, all without any external annotation.
title Retrospective Learning from Interactions
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.13852