Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Krauss, Henrik, Licher, Johann, Takeishi, Naoya, Raatz, Annika, Yairi, Takehisa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914465516093440
author Krauss, Henrik
Licher, Johann
Takeishi, Naoya
Raatz, Annika
Yairi, Takehisa
author_facet Krauss, Henrik
Licher, Johann
Takeishi, Naoya
Raatz, Annika
Yairi, Takehisa
contents Learning soft continuum robot (SCR) dynamics from video offers flexibility but existing methods lack interpretability or rely on prior assumptions. Model-based approaches require prior knowledge and manual design. We bridge this gap by introducing: (1) The Attention Broadcast Decoder (ABCD), a plug-and-play module for autoencoder-based latent dynamics learning that generates pixel-accurate attention maps localizing each latent dimension's contribution while filtering static backgrounds, enabling visual interpretability via spatially grounded latents and on-image overlays. (2) Visual Oscillator Networks (VONs), a 2D latent oscillator network coupled to ABCD attention maps for on-image visualization of learned masses, coupling stiffness, and forces, enabling mechanical interpretability. We validate our approach on single- and double-segment SCRs, demonstrating that ABCD-based models significantly improve multi-step prediction accuracy with 5.8x error reduction for Koopman operators and 3.5x for oscillator networks on a two-segment robot. VONs autonomously discover a chain structure of oscillators. This fully data-driven approach yields compact, mechanically interpretable models with potential relevance for future control applications.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18322
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video
Krauss, Henrik
Licher, Johann
Takeishi, Naoya
Raatz, Annika
Yairi, Takehisa
Robotics
Computer Vision and Pattern Recognition
Machine Learning
Learning soft continuum robot (SCR) dynamics from video offers flexibility but existing methods lack interpretability or rely on prior assumptions. Model-based approaches require prior knowledge and manual design. We bridge this gap by introducing: (1) The Attention Broadcast Decoder (ABCD), a plug-and-play module for autoencoder-based latent dynamics learning that generates pixel-accurate attention maps localizing each latent dimension's contribution while filtering static backgrounds, enabling visual interpretability via spatially grounded latents and on-image overlays. (2) Visual Oscillator Networks (VONs), a 2D latent oscillator network coupled to ABCD attention maps for on-image visualization of learned masses, coupling stiffness, and forces, enabling mechanical interpretability. We validate our approach on single- and double-segment SCRs, demonstrating that ABCD-based models significantly improve multi-step prediction accuracy with 5.8x error reduction for Koopman operators and 3.5x for oscillator networks on a two-segment robot. VONs autonomously discover a chain structure of oscillators. This fully data-driven approach yields compact, mechanically interpretable models with potential relevance for future control applications.
title Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video
topic Robotics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2511.18322