Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914071774756864 |
|---|---|
| author | Jiang, Tianchong Ji, Jingtian Tan, Xiangshan Fang, Jiading Bhattad, Anand Guizilini, Vitor Walter, Matthew R. |
| author_facet | Jiang, Tianchong Ji, Jingtian Tan, Xiangshan Fang, Jiading Bhattad, Anand Guizilini, Vitor Walter, Matthew R. |
| contents | We study view-invariant imitation learning by explicitly conditioning policies on camera extrinsics. Using Plucker embeddings of per-pixel rays, we show that conditioning on extrinsics significantly improves generalization across viewpoints for standard behavior cloning policies, including ACT, Diffusion Policy, and SmolVLA. To evaluate policy robustness under realistic viewpoint shifts, we introduce six manipulation tasks in RoboSuite and ManiSkill that pair "fixed" and "randomized" scene variants, decoupling background cues from camera pose. Our analysis reveals that policies without extrinsics often infer camera pose using visual cues from static backgrounds in fixed scenes; this shortcut collapses when workspace geometry or camera placement shifts. Conditioning on extrinsics restores performance and yields robust RGB-only control without depth. We release the tasks, demonstrations, and code at https://ripl.github.io/know_your_camera/ . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_02268 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning Jiang, Tianchong Ji, Jingtian Tan, Xiangshan Fang, Jiading Bhattad, Anand Guizilini, Vitor Walter, Matthew R. Robotics Computer Vision and Pattern Recognition We study view-invariant imitation learning by explicitly conditioning policies on camera extrinsics. Using Plucker embeddings of per-pixel rays, we show that conditioning on extrinsics significantly improves generalization across viewpoints for standard behavior cloning policies, including ACT, Diffusion Policy, and SmolVLA. To evaluate policy robustness under realistic viewpoint shifts, we introduce six manipulation tasks in RoboSuite and ManiSkill that pair "fixed" and "randomized" scene variants, decoupling background cues from camera pose. Our analysis reveals that policies without extrinsics often infer camera pose using visual cues from static backgrounds in fixed scenes; this shortcut collapses when workspace geometry or camera placement shifts. Conditioning on extrinsics restores performance and yields robust RGB-only control without depth. We release the tasks, demonstrations, and code at https://ripl.github.io/know_your_camera/ . |
| title | Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.02268 |