Articulation in Prime: Primitive-Based Articulated Object Understanding from a Single Casual Video

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Artykov, Arslan, Ravaud, Tom, Violante-Grezzi, Nicolás, Lepetit, Vincent
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917508365156352
author Artykov, Arslan
Ravaud, Tom
Violante-Grezzi, Nicolás
Lepetit, Vincent
author_facet Artykov, Arslan
Ravaud, Tom
Violante-Grezzi, Nicolás
Lepetit, Vincent
contents Retrieving the 3D kinematics of articulated objects from monocular video is a fundamental challenge in computer vision. Existing methods rely on complex video setups or cues such as long-term point tracking or wide-baseline matching, but are frequently brittle under severe occlusions, rapid camera ego-motion, or weak local features. Learning-based methods, meanwhile, struggle to generalize beyond their training categories. We propose a category-agnostic optimization framework that treats articulated object understanding as a primitive-fitting problem. Geometric primitives serve as a proxy representation that avoids the pitfalls of unstable point tracks; a novel mechanism organizes them into coherent parts constrained by revolute and prismatic joints. Our formulation jointly optimizes part segmentation and joint parameters, recovering complex kinematics from a single casually captured video. A visibility-aware procedure handles partial observations and occlusions inherent to real-world data. We also propose the AiP-synth and AiP-real benchmarks, featuring significant camera motion and heavy occlusions, and outperform existing methods. Project page: https://aartykov.github.io/Articulation-in-Prime/
format Preprint
id arxiv_https___arxiv_org_abs_2605_18645
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Articulation in Prime: Primitive-Based Articulated Object Understanding from a Single Casual Video
Artykov, Arslan
Ravaud, Tom
Violante-Grezzi, Nicolás
Lepetit, Vincent
Computer Vision and Pattern Recognition
Retrieving the 3D kinematics of articulated objects from monocular video is a fundamental challenge in computer vision. Existing methods rely on complex video setups or cues such as long-term point tracking or wide-baseline matching, but are frequently brittle under severe occlusions, rapid camera ego-motion, or weak local features. Learning-based methods, meanwhile, struggle to generalize beyond their training categories. We propose a category-agnostic optimization framework that treats articulated object understanding as a primitive-fitting problem. Geometric primitives serve as a proxy representation that avoids the pitfalls of unstable point tracks; a novel mechanism organizes them into coherent parts constrained by revolute and prismatic joints. Our formulation jointly optimizes part segmentation and joint parameters, recovering complex kinematics from a single casually captured video. A visibility-aware procedure handles partial observations and occlusions inherent to real-world data. We also propose the AiP-synth and AiP-real benchmarks, featuring significant camera motion and heavy occlusions, and outperform existing methods. Project page: https://aartykov.github.io/Articulation-in-Prime/
title Articulation in Prime: Primitive-Based Articulated Object Understanding from a Single Casual Video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.18645