Self-Regulation and Requesting Interventions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Min, So Yeon, Wu, Yue, Sun, Jimin, Kaufmann, Max, Tajwar, Fahim, Bisk, Yonatan, Salakhutdinov, Ruslan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916602712162304
author Min, So Yeon
Wu, Yue
Sun, Jimin
Kaufmann, Max
Tajwar, Fahim
Bisk, Yonatan
Salakhutdinov, Ruslan
author_facet Min, So Yeon
Wu, Yue
Sun, Jimin
Kaufmann, Max
Tajwar, Fahim
Bisk, Yonatan
Salakhutdinov, Ruslan
contents Human intelligence involves metacognitive abilities like self-regulation, recognizing limitations, and seeking assistance only when needed. While LLM Agents excel in many domains, they often lack this awareness. Overconfident agents risk catastrophic failures, while those that seek help excessively hinder efficiency. A key challenge is enabling agents with a limited intervention budget $C$ is to decide when to request assistance. In this paper, we propose an offline framework that trains a "helper" policy to request interventions, such as more powerful models or test-time compute, by combining LLM-based process reward models (PRMs) with tabular reinforcement learning. Using state transitions collected offline, we score optimal intervention timing with PRMs and train the helper model on these labeled trajectories. This offline approach significantly reduces costly intervention calls during training. Furthermore, the integration of PRMs with tabular RL enhances robustness to off-policy data while avoiding the inefficiencies of deep RL. We empirically find that our method delivers optimal helper behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04576
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Regulation and Requesting Interventions
Min, So Yeon
Wu, Yue
Sun, Jimin
Kaufmann, Max
Tajwar, Fahim
Bisk, Yonatan
Salakhutdinov, Ruslan
Machine Learning
Computation and Language
Human intelligence involves metacognitive abilities like self-regulation, recognizing limitations, and seeking assistance only when needed. While LLM Agents excel in many domains, they often lack this awareness. Overconfident agents risk catastrophic failures, while those that seek help excessively hinder efficiency. A key challenge is enabling agents with a limited intervention budget $C$ is to decide when to request assistance. In this paper, we propose an offline framework that trains a "helper" policy to request interventions, such as more powerful models or test-time compute, by combining LLM-based process reward models (PRMs) with tabular reinforcement learning. Using state transitions collected offline, we score optimal intervention timing with PRMs and train the helper model on these labeled trajectories. This offline approach significantly reduces costly intervention calls during training. Furthermore, the integration of PRMs with tabular RL enhances robustness to off-policy data while avoiding the inefficiencies of deep RL. We empirically find that our method delivers optimal helper behavior.
title Self-Regulation and Requesting Interventions
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2502.04576