SIRE: SE(3) Intrinsic Rigidity Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Smith, Cameron, Van Hoorick, Basile, Guizilini, Vitor, Wang, Yue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909532951674880
author Smith, Cameron
Van Hoorick, Basile
Guizilini, Vitor
Wang, Yue
author_facet Smith, Cameron
Van Hoorick, Basile
Guizilini, Vitor
Wang, Yue
contents Motion serves as a powerful cue for scene perception and understanding by separating independently moving surfaces and organizing the physical world into distinct entities. We introduce SIRE, a self-supervised method for motion discovery of objects and dynamic scene reconstruction from casual scenes by learning intrinsic rigidity embeddings from videos. Our method trains an image encoder to estimate scene rigidity and geometry, supervised by a simple 4D reconstruction loss: a least-squares solver uses the estimated geometry and rigidity to lift 2D point track trajectories into SE(3) tracks, which are simply re-projected back to 2D and compared against the original 2D trajectories for supervision. Crucially, our framework is fully end-to-end differentiable and can be optimized either on video datasets to learn generalizable image priors, or even on a single video to capture scene-specific structure - highlighting strong data efficiency. We demonstrate the effectiveness of our rigidity embeddings and geometry across multiple settings, including downstream object segmentation, SE(3) rigid motion estimation, and self-supervised depth estimation. Our findings suggest that SIRE can learn strong geometry and motion rigidity priors from video data, with minimal supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2503_07739
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SIRE: SE(3) Intrinsic Rigidity Embeddings
Smith, Cameron
Van Hoorick, Basile
Guizilini, Vitor
Wang, Yue
Computer Vision and Pattern Recognition
Motion serves as a powerful cue for scene perception and understanding by separating independently moving surfaces and organizing the physical world into distinct entities. We introduce SIRE, a self-supervised method for motion discovery of objects and dynamic scene reconstruction from casual scenes by learning intrinsic rigidity embeddings from videos. Our method trains an image encoder to estimate scene rigidity and geometry, supervised by a simple 4D reconstruction loss: a least-squares solver uses the estimated geometry and rigidity to lift 2D point track trajectories into SE(3) tracks, which are simply re-projected back to 2D and compared against the original 2D trajectories for supervision. Crucially, our framework is fully end-to-end differentiable and can be optimized either on video datasets to learn generalizable image priors, or even on a single video to capture scene-specific structure - highlighting strong data efficiency. We demonstrate the effectiveness of our rigidity embeddings and geometry across multiple settings, including downstream object segmentation, SE(3) rigid motion estimation, and self-supervised depth estimation. Our findings suggest that SIRE can learn strong geometry and motion rigidity priors from video data, with minimal supervision.
title SIRE: SE(3) Intrinsic Rigidity Embeddings
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.07739