Orient Anything V2: Unifying Orientation and Rotation Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Zehan, Zhang, Ziang, Xu, Jiayang, Wang, Jialei, Pang, Tianyu, Du, Chao, Zhao, HengShuang, Zhao, Zhou
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917191617609728
author Wang, Zehan
Zhang, Ziang
Xu, Jiayang
Wang, Jialei
Pang, Tianyu
Du, Chao
Zhao, HengShuang
Zhao, Zhou
author_facet Wang, Zehan
Zhang, Ziang
Xu, Jiayang
Wang, Jialei
Pang, Tianyu
Du, Chao
Zhao, HengShuang
Zhao, Zhou
contents This work presents Orient Anything V2, an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. Building upon Orient Anything V1, which defines orientation via a single unique front face, V2 extends this capability to handle objects with diverse rotational symmetries and directly estimate relative rotations. These improvements are enabled by four key innovations: 1) Scalable 3D assets synthesized by generative models, ensuring broad category coverage and balanced data distribution; 2) An efficient, model-in-the-loop annotation system that robustly identifies 0 to N valid front faces for each object; 3) A symmetry-aware, periodic distribution fitting objective that captures all plausible front-facing orientations, effectively modeling object rotational symmetry; 4) A multi-frame architecture that directly predicts relative object rotations. Extensive experiments show that Orient Anything V2 achieves state-of-the-art zero-shot performance on orientation estimation, 6DoF pose estimation, and object symmetry recognition across 11 widely used benchmarks. The model demonstrates strong generalization, significantly broadening the applicability of orientation estimation in diverse downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05573
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Orient Anything V2: Unifying Orientation and Rotation Understanding
Wang, Zehan
Zhang, Ziang
Xu, Jiayang
Wang, Jialei
Pang, Tianyu
Du, Chao
Zhao, HengShuang
Zhao, Zhou
Computer Vision and Pattern Recognition
This work presents Orient Anything V2, an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. Building upon Orient Anything V1, which defines orientation via a single unique front face, V2 extends this capability to handle objects with diverse rotational symmetries and directly estimate relative rotations. These improvements are enabled by four key innovations: 1) Scalable 3D assets synthesized by generative models, ensuring broad category coverage and balanced data distribution; 2) An efficient, model-in-the-loop annotation system that robustly identifies 0 to N valid front faces for each object; 3) A symmetry-aware, periodic distribution fitting objective that captures all plausible front-facing orientations, effectively modeling object rotational symmetry; 4) A multi-frame architecture that directly predicts relative object rotations. Extensive experiments show that Orient Anything V2 achieves state-of-the-art zero-shot performance on orientation estimation, 6DoF pose estimation, and object symmetry recognition across 11 widely used benchmarks. The model demonstrates strong generalization, significantly broadening the applicability of orientation estimation in diverse downstream tasks.
title Orient Anything V2: Unifying Orientation and Rotation Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.05573