SigLoMa: Learning Open-World Quadrupedal Loco-Manipulation from Ego-Centric Vision

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Shiyi, Liu, Haiyi, Yang, Mingye, Zhang, Jiaqi, Zhang, Debing
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914531435872256
author Chen, Shiyi
Liu, Haiyi
Yang, Mingye
Zhang, Jiaqi
Zhang, Debing
author_facet Chen, Shiyi
Liu, Haiyi
Yang, Mingye
Zhang, Jiaqi
Zhang, Debing
contents Designing an open-world quadrupedal loco-manipulation system is highly challenging. Traditional reinforcement learning frameworks utilizing exteroception often suffer from extreme sample inefficiency and massive sim-to-real gaps. Furthermore, the inherent latency of visual tracking fundamentally conflicts with the high-frequency demands of precise floating-base control. Consequently, existing systems lean heavily on expensive external motion capture and off-board computation. To eliminate these dependencies, we present SigLoMa, a fully onboard, ego-centric vision-based pick-and-place framework. At the core of SigLoMa is the introduction of Sigma Points, a lightweight geometric representation for exteroception that guarantees high scalability and native sim-to-real alignment. To bridge the frequency divide between slow perception and fast control, we design an ego-centric Kalman Filter to provide robust, high-rate state estimation. On the learning front, we alleviate sample inefficiency via an Active Sampling Curriculum guided by Hint Poses, and tackle the robot's structural visual blind spots using temporal encoding coupled with simulated random-walk drift. Real-world experiments validate that, relying solely on a 5Hz (200 ms latency) open-vocabulary detector, SigLoMa successfully executes dynamic loco-manipulation across multiple tasks, achieving performance comparable to expert human teleoperation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03846
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SigLoMa: Learning Open-World Quadrupedal Loco-Manipulation from Ego-Centric Vision
Chen, Shiyi
Liu, Haiyi
Yang, Mingye
Zhang, Jiaqi
Zhang, Debing
Robotics
Designing an open-world quadrupedal loco-manipulation system is highly challenging. Traditional reinforcement learning frameworks utilizing exteroception often suffer from extreme sample inefficiency and massive sim-to-real gaps. Furthermore, the inherent latency of visual tracking fundamentally conflicts with the high-frequency demands of precise floating-base control. Consequently, existing systems lean heavily on expensive external motion capture and off-board computation. To eliminate these dependencies, we present SigLoMa, a fully onboard, ego-centric vision-based pick-and-place framework. At the core of SigLoMa is the introduction of Sigma Points, a lightweight geometric representation for exteroception that guarantees high scalability and native sim-to-real alignment. To bridge the frequency divide between slow perception and fast control, we design an ego-centric Kalman Filter to provide robust, high-rate state estimation. On the learning front, we alleviate sample inefficiency via an Active Sampling Curriculum guided by Hint Poses, and tackle the robot's structural visual blind spots using temporal encoding coupled with simulated random-walk drift. Real-world experiments validate that, relying solely on a 5Hz (200 ms latency) open-vocabulary detector, SigLoMa successfully executes dynamic loco-manipulation across multiple tasks, achieving performance comparable to expert human teleoperation.
title SigLoMa: Learning Open-World Quadrupedal Loco-Manipulation from Ego-Centric Vision
topic Robotics
url https://arxiv.org/abs/2605.03846