Orient Anything V2: Unifying Orientation and Rotation Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zehan, Zhang, Ziang, Xu, Jiayang, Wang, Jialei, Pang, Tianyu, Du, Chao, Zhao, HengShuang, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
GenSpace: Benchmarking Spatially-Aware Image Generation
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
Depth Anything with Any Prior
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
by: Zhou, Junbao, et al.
Published: (2025)
by: Zhou, Junbao, et al.
Published: (2025)
DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
by: Zhang, Ziang, et al.
Published: (2025)
by: Zhang, Ziang, et al.
Published: (2025)
Depth Anything V2
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
by: Huang, Haifeng, et al.
Published: (2026)
by: Huang, Haifeng, et al.
Published: (2026)
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
RQFormer: Rotated Query Transformer for End-to-End Oriented Object Detection
by: Zhao, Jiaqi, et al.
Published: (2023)
by: Zhao, Jiaqi, et al.
Published: (2023)
ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks
by: Xu, Jiayang, et al.
Published: (2026)
by: Xu, Jiayang, et al.
Published: (2026)
MESA: Matching Everything by Segmenting Anything
by: Zhang, Yesheng, et al.
Published: (2024)
by: Zhang, Yesheng, et al.
Published: (2024)
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design
by: Schneider, Benjamin, et al.
Published: (2025)
by: Schneider, Benjamin, et al.
Published: (2025)
Unsafe by Reciprocity: How Generation-Understanding Coupling Undermines Safety in Unified Multimodal Models
by: Wang, Kaishen, et al.
Published: (2026)
by: Wang, Kaishen, et al.
Published: (2026)
SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification
by: Zhang, Jiacheng, et al.
Published: (2026)
by: Zhang, Jiacheng, et al.
Published: (2026)
PointSAM: Pointly-Supervised Segment Anything Model for Remote Sensing Images
by: Liu, Nanqing, et al.
Published: (2024)
by: Liu, Nanqing, et al.
Published: (2024)
Describe Anything in Medical Images
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic Segmentation
by: Zhong, Ding, et al.
Published: (2025)
by: Zhong, Ding, et al.
Published: (2025)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
by: Huang, Haifeng, et al.
Published: (2023)
by: Huang, Haifeng, et al.
Published: (2023)
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
by: Hong, Minjie, et al.
Published: (2025)
by: Hong, Minjie, et al.
Published: (2025)
Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
by: Du, Henghui, et al.
Published: (2025)
by: Du, Henghui, et al.
Published: (2025)
MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos
by: Gong, Kehong, et al.
Published: (2025)
by: Gong, Kehong, et al.
Published: (2025)
Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Stereo Anything: Unifying Zero-shot Stereo Matching with Large-Scale Mixed Data
by: Guo, Xianda, et al.
Published: (2024)
by: Guo, Xianda, et al.
Published: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024)
by: Cheng, Xize, et al.
Published: (2024)
Prompting Segment Anything Model with Domain-Adaptive Prototype for Generalizable Medical Image Segmentation
by: Wei, Zhikai, et al.
Published: (2024)
by: Wei, Zhikai, et al.
Published: (2024)
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
by: Xu, Linrui, et al.
Published: (2024)
by: Xu, Linrui, et al.
Published: (2024)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
by: Zhang, Ruiheng, et al.
Published: (2026)
by: Zhang, Ruiheng, et al.
Published: (2026)
GRA: Detecting Oriented Objects through Group-wise Rotating and Attention
by: Wang, Jiangshan, et al.
Published: (2024)
by: Wang, Jiangshan, et al.
Published: (2024)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
Training for Identity, Inference for Controllability: A Unified Approach to Tuning-Free Face Personalization
by: Pang, Lianyu, et al.
Published: (2025)
by: Pang, Lianyu, et al.
Published: (2025)
OnlineX: Unified Online 3D Reconstruction and Understanding with Active-to-Stable State Evolution
by: Xia, Chong, et al.
Published: (2026)
by: Xia, Chong, et al.
Published: (2026)
Contour Field based Elliptical Shape Prior for the Segment Anything Model
by: Zhao, Xinyu, et al.
Published: (2025)
by: Zhao, Xinyu, et al.
Published: (2025)
Uni-Sign: Toward Unified Sign Language Understanding at Scale
by: Li, Zecheng, et al.
Published: (2025)
by: Li, Zecheng, et al.
Published: (2025)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
FastDrag: Manipulate Anything in One Step
by: Zhao, Xuanjia, et al.
Published: (2024)
by: Zhao, Xuanjia, et al.
Published: (2024)
Count Anything
by: Lei, Mengqi, et al.
Published: (2026)
by: Lei, Mengqi, et al.
Published: (2026)
Similar Items
-
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
by: Wang, Zehan, et al.
Published: (2024) -
GenSpace: Benchmarking Spatially-Aware Image Generation
by: Wang, Zehan, et al.
Published: (2025) -
Depth Anything with Any Prior
by: Wang, Zehan, et al.
Published: (2025) -
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024) -
Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
by: Zhou, Junbao, et al.
Published: (2025)