Sim-Anchored Learning for On-the-Fly Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mabsout, Bassel El, Roozkhosh, Shahin, Mysore, Siddharth, Saenko, Kate, Mancuso, Renato
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908344441110528
author Mabsout, Bassel El
Roozkhosh, Shahin
Mysore, Siddharth
Saenko, Kate
Mancuso, Renato
author_facet Mabsout, Bassel El
Roozkhosh, Shahin
Mysore, Siddharth
Saenko, Kate
Mancuso, Renato
contents Fine-tuning simulation-trained RL agents with real-world data often degrades crucial behaviors due to limited or skewed data distributions. We argue that designer priorities exist not just in reward functions, but also in simulation design choices like task selection and state initialization. When adapting to real-world data, agents can experience catastrophic forgetting in important but underrepresented scenarios. We propose framing live-adaptation as a multi-objective optimization problem, where policy objectives must be satisfied both in simulation and reality. Our approach leverages critics from simulation as "anchors for design intent" (anchor critics). By jointly optimizing policies against both anchor critics and critics trained on real-world experience, our method enables adaptation while preserving prioritized behaviors from simulation. Evaluations demonstrate robust behavior retention in sim-to-sim benchmarks and a sim-to-real scenario with a racing quadrotor, allowing for power consumption reductions of up to 50% without control loss. We also contribute SwaNNFlight, an open-source firmware for enabling live adaptation on similar robotic platforms.
format Preprint
id arxiv_https___arxiv_org_abs_2301_06987
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Sim-Anchored Learning for On-the-Fly Adaptation
Mabsout, Bassel El
Roozkhosh, Shahin
Mysore, Siddharth
Saenko, Kate
Mancuso, Renato
Robotics
Machine Learning
Fine-tuning simulation-trained RL agents with real-world data often degrades crucial behaviors due to limited or skewed data distributions. We argue that designer priorities exist not just in reward functions, but also in simulation design choices like task selection and state initialization. When adapting to real-world data, agents can experience catastrophic forgetting in important but underrepresented scenarios. We propose framing live-adaptation as a multi-objective optimization problem, where policy objectives must be satisfied both in simulation and reality. Our approach leverages critics from simulation as "anchors for design intent" (anchor critics). By jointly optimizing policies against both anchor critics and critics trained on real-world experience, our method enables adaptation while preserving prioritized behaviors from simulation. Evaluations demonstrate robust behavior retention in sim-to-sim benchmarks and a sim-to-real scenario with a racing quadrotor, allowing for power consumption reductions of up to 50% without control loss. We also contribute SwaNNFlight, an open-source firmware for enabling live adaptation on similar robotic platforms.
title Sim-Anchored Learning for On-the-Fly Adaptation
topic Robotics
Machine Learning
url https://arxiv.org/abs/2301.06987