AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xuying, Zhou, Yupeng, Wang, Kai, Wang, Yikai, Li, Zhen, Jiao, Shaohui, Zhou, Daquan, Hou, Qibin, Cheng, Ming-Ming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024)
by: Zhou, Yupeng, et al.
Published: (2024)
GeoWorld: Unlocking the Potential of Geometry Models to Facilitate High-Fidelity 3D Scene Generation
by: Wan, Yuhao, et al.
Published: (2025)
by: Wan, Yuhao, et al.
Published: (2025)
Sora Generates Videos with Stunning Geometrical Consistency
by: Li, Xuanyi, et al.
Published: (2024)
by: Li, Xuanyi, et al.
Published: (2024)
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction
by: Zhang, Xuying, et al.
Published: (2024)
by: Zhang, Xuying, et al.
Published: (2024)
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
by: Zhou, Yupeng, et al.
Published: (2026)
by: Zhou, Yupeng, et al.
Published: (2026)
SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-Resolution
by: Zhou, Yupeng, et al.
Published: (2023)
by: Zhou, Yupeng, et al.
Published: (2023)
Towards Stable 3D Object Detection
by: Wang, Jiabao, et al.
Published: (2024)
by: Wang, Jiabao, et al.
Published: (2024)
Referring Camouflaged Object Detection
by: Zhang, Xuying, et al.
Published: (2023)
by: Zhang, Xuying, et al.
Published: (2023)
High-Quality Mask Tuning Matters for Open-Vocabulary Segmentation
by: Zeng, Quan-Sheng, et al.
Published: (2024)
by: Zeng, Quan-Sheng, et al.
Published: (2024)
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
by: Zhou, Yupeng, et al.
Published: (2025)
by: Zhou, Yupeng, et al.
Published: (2025)
FlexiDreamer: Single Image-to-3D Generation with FlexiCubes
by: Zhao, Ruowen, et al.
Published: (2024)
by: Zhao, Ruowen, et al.
Published: (2024)
Dehallu3D: Hallucination-Mitigated 3D Generation from Single Image via Cyclic View Consistency Refinement
by: Wang, Xiwen, et al.
Published: (2026)
by: Wang, Xiwen, et al.
Published: (2026)
OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
by: Yin, Bo-Wen, et al.
Published: (2025)
by: Yin, Bo-Wen, et al.
Published: (2025)
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
by: Li, Yunheng, et al.
Published: (2025)
by: Li, Yunheng, et al.
Published: (2025)
DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
by: Yin, Bowen, et al.
Published: (2023)
by: Yin, Bowen, et al.
Published: (2023)
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
by: Ouyang, Ziheng, et al.
Published: (2025)
by: Ouyang, Ziheng, et al.
Published: (2025)
Zone Evaluation: Revealing Spatial Bias in Object Detection
by: Zheng, Zhaohui, et al.
Published: (2023)
by: Zheng, Zhaohui, et al.
Published: (2023)
HawkI: Homography & Mutual Information Guidance for 3D-free Single Image to Aerial View
by: Kothandaraman, Divya, et al.
Published: (2023)
by: Kothandaraman, Divya, et al.
Published: (2023)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
by: Wang, Jiabao, et al.
Published: (2023)
by: Wang, Jiabao, et al.
Published: (2023)
ControlSR: Taming Diffusion Models for Consistent Real-World Image Super Resolution
by: Wan, Yuhao, et al.
Published: (2024)
by: Wan, Yuhao, et al.
Published: (2024)
Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions
by: Li, Yunheng, et al.
Published: (2026)
by: Li, Yunheng, et al.
Published: (2026)
DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
by: Yin, Bo-Wen, et al.
Published: (2025)
by: Yin, Bo-Wen, et al.
Published: (2025)
Isotropic3D: Image-to-3D Generation Based on a Single CLIP Embedding
by: Liu, Pengkun, et al.
Published: (2024)
by: Liu, Pengkun, et al.
Published: (2024)
Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
ViewCraft3D: High-Fidelity and View-Consistent 3D Vector Graphics Synthesis
by: Wang, Chuang, et al.
Published: (2025)
by: Wang, Chuang, et al.
Published: (2025)
Mixture of Style Experts for Diverse Image Stylization
by: Zhu, Shihao, et al.
Published: (2026)
by: Zhu, Shihao, et al.
Published: (2026)
GraphicsDreamer: Image to 3D Generation with Physical Consistency
by: Chen, Pei, et al.
Published: (2024)
by: Chen, Pei, et al.
Published: (2024)
Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing
by: Hu, Taihang, et al.
Published: (2025)
by: Hu, Taihang, et al.
Published: (2025)
Human Video Generation from a Single Image with 3D Pose and View Control
by: Wang, Tiantian, et al.
Published: (2026)
by: Wang, Tiantian, et al.
Published: (2026)
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
by: Chen, Yuming, et al.
Published: (2023)
by: Chen, Yuming, et al.
Published: (2023)
AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation
by: Liang, Xinyue, et al.
Published: (2025)
by: Liang, Xinyue, et al.
Published: (2025)
DecoRec: Decomposed 3D Scene Reconstruction from Single-View Images via Object-Level Diffusion
by: Ping, Yuhan, et al.
Published: (2026)
by: Ping, Yuhan, et al.
Published: (2026)
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
by: Lu, Keyang, et al.
Published: (2025)
by: Lu, Keyang, et al.
Published: (2025)
LucidDreaming: Controllable Object-Centric 3D Generation
by: Wang, Zhaoning, et al.
Published: (2023)
by: Wang, Zhaoning, et al.
Published: (2023)
DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion
by: Qu, Yansong, et al.
Published: (2025)
by: Qu, Yansong, et al.
Published: (2025)
Interact3D: Compositional 3D Generation of Interactive Objects
by: Shan, Hui, et al.
Published: (2026)
by: Shan, Hui, et al.
Published: (2026)
PB-NBV: Efficient Projection-Based Next-Best-View Planning Framework for Reconstruction of Unknown Objects
by: Jia, Zhizhou, et al.
Published: (2025)
by: Jia, Zhizhou, et al.
Published: (2025)
Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation
by: Lin, Siyou, et al.
Published: (2026)
by: Lin, Siyou, et al.
Published: (2026)
Similar Items
-
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024) -
GeoWorld: Unlocking the Potential of Geometry Models to Facilitate High-Fidelity 3D Scene Generation
by: Wan, Yuhao, et al.
Published: (2025) -
Sora Generates Videos with Stunning Geometrical Consistency
by: Li, Xuanyi, et al.
Published: (2024) -
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction
by: Zhang, Xuying, et al.
Published: (2024) -
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
by: Zhou, Yupeng, et al.
Published: (2026)