Adaptive Policy Backbone via Shared Network

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Bumgeun, Lee, Donghwan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909809217896448
author Park, Bumgeun
Lee, Donghwan
author_facet Park, Bumgeun
Lee, Donghwan
contents Reinforcement learning (RL) has achieved impressive results across domains, yet learning an optimal policy typically requires extensive interaction data, limiting practical deployment. A common remedy is to leverage priors, such as pre-collected datasets or reference policies, but their utility degrades under task mismatch between training and deployment. While prior work has sought to address this mismatch, it has largely been restricted to in-distribution settings. To address this challenge, we propose Adaptive Policy Backbone (APB), a meta-transfer RL method that inserts lightweight linear layers before and after a shared backbone, thereby enabling parameter-efficient fine-tuning (PEFT) while preserving prior knowledge during adaptation. Our results show that APB improves sample efficiency over standard RL and adapts to out-of-distribution (OOD) tasks where existing meta-RL baselines typically fail.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22310
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptive Policy Backbone via Shared Network
Park, Bumgeun
Lee, Donghwan
Machine Learning
Artificial Intelligence
Reinforcement learning (RL) has achieved impressive results across domains, yet learning an optimal policy typically requires extensive interaction data, limiting practical deployment. A common remedy is to leverage priors, such as pre-collected datasets or reference policies, but their utility degrades under task mismatch between training and deployment. While prior work has sought to address this mismatch, it has largely been restricted to in-distribution settings. To address this challenge, we propose Adaptive Policy Backbone (APB), a meta-transfer RL method that inserts lightweight linear layers before and after a shared backbone, thereby enabling parameter-efficient fine-tuning (PEFT) while preserving prior knowledge during adaptation. Our results show that APB improves sample efficiency over standard RL and adapts to out-of-distribution (OOD) tasks where existing meta-RL baselines typically fail.
title Adaptive Policy Backbone via Shared Network
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.22310