Object-Centric Latent Action Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915740994502656 |
|---|---|
| author | Klepach, Albina Nikulin, Alexander Zisman, Ilya Tarasov, Denis Derevyagin, Alexander Polubarov, Andrei Lyubaykin, Nikita Kiselev, Igor Kurenkov, Vladislav |
| author_facet | Klepach, Albina Nikulin, Alexander Zisman, Ilya Tarasov, Denis Derevyagin, Alexander Polubarov, Andrei Lyubaykin, Nikita Kiselev, Igor Kurenkov, Vladislav |
| contents | Leveraging vast amounts of unlabeled internet video data for embodied AI is currently bottlenecked by the lack of action labels and the presence of action-correlated visual distractors. Although recent latent action policy optimization (LAPO) has shown promise in inferring proxy action labels from visual observations, its performance degrades significantly when distractors are present. To address this limitation, we propose a novel object-centric latent action learning framework that centers on objects rather than pixels. We leverage self-supervised object-centric pretraining to disentangle the movement of the agent and distracting background dynamics. This allows LAPO to focus on task-relevant interactions, resulting in more robust proxy-action labels, enabling better imitation learning and efficient adaptation of the agent with just a few action-labeled trajectories. We evaluated our method in eight visually complex tasks across the Distracting Control Suite (DCS) and Distracting MetaWorld (DMW). Our results show that object-centric pretraining mitigates the negative effects of distractors by 50%, as measured by downstream task performance: average return (DCS) and success rate (DMW). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_09680 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Object-Centric Latent Action Learning Klepach, Albina Nikulin, Alexander Zisman, Ilya Tarasov, Denis Derevyagin, Alexander Polubarov, Andrei Lyubaykin, Nikita Kiselev, Igor Kurenkov, Vladislav Computer Vision and Pattern Recognition Artificial Intelligence Leveraging vast amounts of unlabeled internet video data for embodied AI is currently bottlenecked by the lack of action labels and the presence of action-correlated visual distractors. Although recent latent action policy optimization (LAPO) has shown promise in inferring proxy action labels from visual observations, its performance degrades significantly when distractors are present. To address this limitation, we propose a novel object-centric latent action learning framework that centers on objects rather than pixels. We leverage self-supervised object-centric pretraining to disentangle the movement of the agent and distracting background dynamics. This allows LAPO to focus on task-relevant interactions, resulting in more robust proxy-action labels, enabling better imitation learning and efficient adaptation of the agent with just a few action-labeled trajectories. We evaluated our method in eight visually complex tasks across the Distracting Control Suite (DCS) and Distracting MetaWorld (DMW). Our results show that object-centric pretraining mitigates the negative effects of distractors by 50%, as measured by downstream task performance: average return (DCS) and success rate (DMW). |
| title | Object-Centric Latent Action Learning |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2502.09680 |