Multi-Task Reinforcement Learning with Language-Encoded Gated Policy Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Arora, Rushiv
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909829977604096
author Arora, Rushiv
author_facet Arora, Rushiv
contents Multi-task reinforcement learning often relies on task metadata -- such as brief natural-language descriptions -- to guide behavior across diverse objectives. We present Lexical Policy Networks (LEXPOL), a language-conditioned mixture-of-policies architecture for multi-task RL. LEXPOL encodes task metadata with a text encoder and uses a learned gating module to select or blend among multiple sub-policies, enabling end-to-end training across tasks. On MetaWorld benchmarks, LEXPOL matches or exceeds strong multi-task baselines in success rate and sample efficiency, without task-specific retraining. To analyze the mechanism, we further study settings with fixed expert policies obtained independently of the gate and show that the learned language gate composes these experts to produce behaviors appropriate to novel task descriptions and unseen task combinations. These results indicate that natural-language metadata can effectively index and recombine reusable skills within a single policy.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06138
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Task Reinforcement Learning with Language-Encoded Gated Policy Networks
Arora, Rushiv
Machine Learning
Artificial Intelligence
I.2.6
Multi-task reinforcement learning often relies on task metadata -- such as brief natural-language descriptions -- to guide behavior across diverse objectives. We present Lexical Policy Networks (LEXPOL), a language-conditioned mixture-of-policies architecture for multi-task RL. LEXPOL encodes task metadata with a text encoder and uses a learned gating module to select or blend among multiple sub-policies, enabling end-to-end training across tasks. On MetaWorld benchmarks, LEXPOL matches or exceeds strong multi-task baselines in success rate and sample efficiency, without task-specific retraining. To analyze the mechanism, we further study settings with fixed expert policies obtained independently of the gate and show that the learned language gate composes these experts to produce behaviors appropriate to novel task descriptions and unseen task combinations. These results indicate that natural-language metadata can effectively index and recombine reusable skills within a single policy.
title Multi-Task Reinforcement Learning with Language-Encoded Gated Policy Networks
topic Machine Learning
Artificial Intelligence
I.2.6
url https://arxiv.org/abs/2510.06138