$π^3$: Permutation-Equivariant Visual Geometry Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Yifan, Zhou, Jianjun, Zhu, Haoyi, Chang, Wenzheng, Zhou, Yang, Li, Zizun, Chen, Junyi, Pang, Jiangmiao, Shen, Chunhua, He, Tong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910044415590400
author Wang, Yifan
Zhou, Jianjun
Zhu, Haoyi
Chang, Wenzheng
Zhou, Yang
Li, Zizun
Chen, Junyi
Pang, Jiangmiao
Shen, Chunhua
He, Tong
author_facet Wang, Yifan
Zhou, Jianjun
Zhu, Haoyi
Chang, Wenzheng
Zhou, Yang
Li, Zizun
Chen, Junyi
Pang, Jiangmiao
Shen, Chunhua
He, Tong
contents We introduce $π^3$, a feed-forward neural network that offers a novel approach to visual geometry reconstruction, breaking the reliance on a conventional fixed reference view. Previous methods often anchor their reconstructions to a designated viewpoint, an inductive bias that can lead to instability and failures if the reference is suboptimal. In contrast, $π^3$ employs a fully permutation-equivariant architecture to predict affine-invariant camera poses and scale-invariant local point maps without any reference frames. This design not only makes our model inherently robust to input ordering, but also leads to higher accuracy and performance. These advantages enable our simple and bias-free approach to achieve state-of-the-art performance on a wide range of tasks, including camera pose estimation, monocular/video depth estimation, and dense point map reconstruction. Code and models are available at https://github.com/yyfz/Pi3.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13347
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle $π^3$: Permutation-Equivariant Visual Geometry Learning
Wang, Yifan
Zhou, Jianjun
Zhu, Haoyi
Chang, Wenzheng
Zhou, Yang
Li, Zizun
Chen, Junyi
Pang, Jiangmiao
Shen, Chunhua
He, Tong
Computer Vision and Pattern Recognition
We introduce $π^3$, a feed-forward neural network that offers a novel approach to visual geometry reconstruction, breaking the reliance on a conventional fixed reference view. Previous methods often anchor their reconstructions to a designated viewpoint, an inductive bias that can lead to instability and failures if the reference is suboptimal. In contrast, $π^3$ employs a fully permutation-equivariant architecture to predict affine-invariant camera poses and scale-invariant local point maps without any reference frames. This design not only makes our model inherently robust to input ordering, but also leads to higher accuracy and performance. These advantages enable our simple and bias-free approach to achieve state-of-the-art performance on a wide range of tasks, including camera pose estimation, monocular/video depth estimation, and dense point map reconstruction. Code and models are available at https://github.com/yyfz/Pi3.
title $π^3$: Permutation-Equivariant Visual Geometry Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.13347