GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Hanxin, Wang, Cong, Tu, Peiyan, Luo, Jiayi, He, Tianyu, Jin, Xin, Chen, Zhibo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Embody4D: A Generalist 4D World Model for Embodied AI
by: Tu, Peiyan, et al.
Published: (2026)
by: Tu, Peiyan, et al.
Published: (2026)
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
by: Wang, Cong, et al.
Published: (2026)
by: Wang, Cong, et al.
Published: (2026)
GaussianSR: 3D Gaussian Super-Resolution with 2D Diffusion Priors
by: Yu, Xiqian, et al.
Published: (2024)
by: Yu, Xiqian, et al.
Published: (2024)
Compositional 3D-aware Video Generation with LLM Director
by: Zhu, Hanxin, et al.
Published: (2024)
by: Zhu, Hanxin, et al.
Published: (2024)
CMC: Few-shot Novel View Synthesis via Cross-view Multiplane Consistency
by: Zhu, Hanxin, et al.
Published: (2024)
by: Zhu, Hanxin, et al.
Published: (2024)
4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
AR4D: Autoregressive 4D Generation from Monocular Videos
by: Zhu, Hanxin, et al.
Published: (2025)
by: Zhu, Hanxin, et al.
Published: (2025)
TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
Is Vanilla MLP in Neural Radiance Field Enough for Few-shot View Synthesis?
by: Zhu, Hanxin, et al.
Published: (2024)
by: Zhu, Hanxin, et al.
Published: (2024)
P-4DGS: Predictive 4D Gaussian Splatting with 90$\times$ Compression
by: Wang, Henan, et al.
Published: (2025)
by: Wang, Henan, et al.
Published: (2025)
Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
by: Luo, Jiayi, et al.
Published: (2026)
by: Luo, Jiayi, et al.
Published: (2026)
End-to-End Rate-Distortion Optimized 3D Gaussian Representation
by: Wang, Henan, et al.
Published: (2024)
by: Wang, Henan, et al.
Published: (2024)
GSemSplat: Generalizable Semantic 3D Gaussian Splatting from Uncalibrated Image Pairs
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
SeD: Semantic-Aware Discriminator for Image Super-Resolution
by: Li, Bingchen, et al.
Published: (2024)
by: Li, Bingchen, et al.
Published: (2024)
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
by: Wu, Haoyu, et al.
Published: (2025)
by: Wu, Haoyu, et al.
Published: (2025)
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
by: Luo, Jiayi, et al.
Published: (2026)
by: Luo, Jiayi, et al.
Published: (2026)
Orchid: Image Latent Diffusion for Joint Appearance and Geometry Generation
by: Krishnan, Akshay, et al.
Published: (2025)
by: Krishnan, Akshay, et al.
Published: (2025)
Light Field Compression Based on Implicit Neural Representation
by: Wang, Henan, et al.
Published: (2024)
by: Wang, Henan, et al.
Published: (2024)
UCIP: A Universal Framework for Compressed Image Super-Resolution using Dynamic Prompt
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration
by: Xie, Siyi, et al.
Published: (2025)
by: Xie, Siyi, et al.
Published: (2025)
3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation
by: He, Yunhong, et al.
Published: (2025)
by: He, Yunhong, et al.
Published: (2025)
Layout2Scene: 3D Semantic Layout Guided Scene Generation via Geometry and Appearance Diffusion Priors
by: Chen, Minglin, et al.
Published: (2025)
by: Chen, Minglin, et al.
Published: (2025)
TAPESTRY: From Geometry to Appearance via Consistent Turntable Videos
by: Zeng, Yan, et al.
Published: (2026)
by: Zeng, Yan, et al.
Published: (2026)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
by: Sun, Jingwen, et al.
Published: (2026)
by: Sun, Jingwen, et al.
Published: (2026)
PEGAsus: 3D Personalization of Geometry and Appearance
by: Hu, Jingyu, et al.
Published: (2026)
by: Hu, Jingyu, et al.
Published: (2026)
EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision
by: Cao, Yifei, et al.
Published: (2025)
by: Cao, Yifei, et al.
Published: (2025)
WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion
by: Kong, Hanyang, et al.
Published: (2025)
by: Kong, Hanyang, et al.
Published: (2025)
DiGA3D: Coarse-to-Fine Diffusional Propagation of Geometry and Appearance for Versatile 3D Inpainting
by: Pan, Jingyi, et al.
Published: (2025)
by: Pan, Jingyi, et al.
Published: (2025)
LiftVSR: Lifting Image Diffusion to Video Super-Resolution via Hybrid Temporal Modeling with Only 4$\times$RTX 4090s
by: Wang, Xijun, et al.
Published: (2025)
by: Wang, Xijun, et al.
Published: (2025)
Learning to Align Generative Appearance Priors for Fine-grained Image Retrieval
by: Wang, Shijie, et al.
Published: (2026)
by: Wang, Shijie, et al.
Published: (2026)
GenTron: Diffusion Transformers for Image and Video Generation
by: Chen, Shoufa, et al.
Published: (2023)
by: Chen, Shoufa, et al.
Published: (2023)
IPDreamer: Appearance-Controllable 3D Object Generation with Complex Image Prompts
by: Zeng, Bohan, et al.
Published: (2023)
by: Zeng, Bohan, et al.
Published: (2023)
GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
by: Miyato, Takeru, et al.
Published: (2023)
by: Miyato, Takeru, et al.
Published: (2023)
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
by: Xing, Ke, et al.
Published: (2025)
by: Xing, Ke, et al.
Published: (2025)
Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery
by: Kang, Yijie, et al.
Published: (2025)
by: Kang, Yijie, et al.
Published: (2025)
Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
by: Huang, Tianyu, et al.
Published: (2025)
by: Huang, Tianyu, et al.
Published: (2025)
FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction
by: Dai, Yixiang, et al.
Published: (2025)
by: Dai, Yixiang, et al.
Published: (2025)
CoNo: Consistency Noise Injection for Tuning-free Long Video Diffusion
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
UniLat3D: Geometry-Appearance Unified Latents for Single-Stage 3D Generation
by: Wu, Guanjun, et al.
Published: (2025)
by: Wu, Guanjun, et al.
Published: (2025)
Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
by: Kim, Jiwon, et al.
Published: (2025)
by: Kim, Jiwon, et al.
Published: (2025)
Similar Items
-
Embody4D: A Generalist 4D World Model for Embodied AI
by: Tu, Peiyan, et al.
Published: (2026) -
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
by: Wang, Cong, et al.
Published: (2026) -
GaussianSR: 3D Gaussian Super-Resolution with 2D Diffusion Priors
by: Yu, Xiqian, et al.
Published: (2024) -
Compositional 3D-aware Video Generation with LLM Director
by: Zhu, Hanxin, et al.
Published: (2024) -
CMC: Few-shot Novel View Synthesis via Cross-view Multiplane Consistency
by: Zhu, Hanxin, et al.
Published: (2024)