GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Mengchen, Wu, Tong, Tan, Jing, Liu, Ziwei, Wetzstein, Gordon, Lin, Dahua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
di: Yang, Shuai, et al.
Pubblicazione: (2024)
di: Yang, Shuai, et al.
Pubblicazione: (2024)
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
di: Zhang, Yuhan, et al.
Pubblicazione: (2025)
di: Zhang, Yuhan, et al.
Pubblicazione: (2025)
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
di: Wu, Tong, et al.
Pubblicazione: (2024)
di: Wu, Tong, et al.
Pubblicazione: (2024)
Video World Models with Long-term Spatial Memory
di: Wu, Tong, et al.
Pubblicazione: (2025)
di: Wu, Tong, et al.
Pubblicazione: (2025)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
di: Wu, Tong, et al.
Pubblicazione: (2024)
di: Wu, Tong, et al.
Pubblicazione: (2024)
SS4D: Native 4D Generative Model via Structured Spacetime Latents
di: Li, Zhibing, et al.
Pubblicazione: (2025)
di: Li, Zhibing, et al.
Pubblicazione: (2025)
IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations
di: Li, Zhibing, et al.
Pubblicazione: (2024)
di: Li, Zhibing, et al.
Pubblicazione: (2024)
Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
di: Fang, Ye, et al.
Pubblicazione: (2024)
di: Fang, Ye, et al.
Pubblicazione: (2024)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
di: Zhang, Mengchen, et al.
Pubblicazione: (2025)
di: Zhang, Mengchen, et al.
Pubblicazione: (2025)
Imagine360: Immersive 360 Video Generation from Perspective Anchor
di: Tan, Jing, et al.
Pubblicazione: (2024)
di: Tan, Jing, et al.
Pubblicazione: (2024)
Omni6D: Large-Vocabulary 3D Object Dataset for Category-Level 6D Object Pose Estimation
di: Zhang, Mengchen, et al.
Pubblicazione: (2024)
di: Zhang, Mengchen, et al.
Pubblicazione: (2024)
Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control
di: Xie, Linxi, et al.
Pubblicazione: (2026)
di: Xie, Linxi, et al.
Pubblicazione: (2026)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
di: He, Hao, et al.
Pubblicazione: (2024)
di: He, Hao, et al.
Pubblicazione: (2024)
BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
di: Wang, Yiming, et al.
Pubblicazione: (2025)
di: Wang, Yiming, et al.
Pubblicazione: (2025)
Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
di: Ackermann, Jan, et al.
Pubblicazione: (2025)
di: Ackermann, Jan, et al.
Pubblicazione: (2025)
Temporal-Mapping Photography for Event Cameras
di: Bao, Yuhan, et al.
Pubblicazione: (2024)
di: Bao, Yuhan, et al.
Pubblicazione: (2024)
TC4D: Trajectory-Conditioned Text-to-4D Generation
di: Bahmani, Sherwin, et al.
Pubblicazione: (2024)
di: Bahmani, Sherwin, et al.
Pubblicazione: (2024)
CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
di: Yang, Qirui, et al.
Pubblicazione: (2025)
di: Yang, Qirui, et al.
Pubblicazione: (2025)
Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images
di: Qi, Zhangyang, et al.
Pubblicazione: (2024)
di: Qi, Zhangyang, et al.
Pubblicazione: (2024)
Neural Ganglion Sensors: Learning Task-specific Event Cameras Inspired by the Neural Circuit of the Human Retina
di: So, Haley M., et al.
Pubblicazione: (2025)
di: So, Haley M., et al.
Pubblicazione: (2025)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
di: Kuang, Zhengfei, et al.
Pubblicazione: (2024)
di: Kuang, Zhengfei, et al.
Pubblicazione: (2024)
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
di: Yuan, Yu, et al.
Pubblicazione: (2024)
di: Yuan, Yu, et al.
Pubblicazione: (2024)
Time-Aware Auto White Balance in Mobile Photography
di: Afifi, Mahmoud, et al.
Pubblicazione: (2025)
di: Afifi, Mahmoud, et al.
Pubblicazione: (2025)
EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh Generation
di: Tang, Jiaxiang, et al.
Pubblicazione: (2024)
di: Tang, Jiaxiang, et al.
Pubblicazione: (2024)
Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
di: Zhang, Yuhan, et al.
Pubblicazione: (2025)
di: Zhang, Yuhan, et al.
Pubblicazione: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
di: Fang, Ye, et al.
Pubblicazione: (2025)
di: Fang, Ye, et al.
Pubblicazione: (2025)
Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text
di: Li, Xinyang, et al.
Pubblicazione: (2024)
di: Li, Xinyang, et al.
Pubblicazione: (2024)
CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models
di: He, Hao, et al.
Pubblicazione: (2025)
di: He, Hao, et al.
Pubblicazione: (2025)
Stencil: Subject-Driven Generation with Context Guidance
di: Chen, Gordon, et al.
Pubblicazione: (2025)
di: Chen, Gordon, et al.
Pubblicazione: (2025)
Infinite Gaze Generation for Videos with Autoregressive Diffusion
di: Kang, Jenna, et al.
Pubblicazione: (2026)
di: Kang, Jenna, et al.
Pubblicazione: (2026)
AutoDirector: Online Auto-scheduling Agents for Multi-sensory Composition
di: Ni, Minheng, et al.
Pubblicazione: (2024)
di: Ni, Minheng, et al.
Pubblicazione: (2024)
Spectral Progressive Diffusion for Efficient Image and Video Generation
di: Xiao, Howard, et al.
Pubblicazione: (2026)
di: Xiao, Howard, et al.
Pubblicazione: (2026)
Foveated Diffusion: Efficient Spatially Adaptive Image and Video Generation
di: Chao, Brian, et al.
Pubblicazione: (2026)
di: Chao, Brian, et al.
Pubblicazione: (2026)
FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views
di: Zhang, Shangzhan, et al.
Pubblicazione: (2025)
di: Zhang, Shangzhan, et al.
Pubblicazione: (2025)
GazeFusion: Saliency-Guided Image Generation
di: Zhang, Yunxiang, et al.
Pubblicazione: (2024)
di: Zhang, Yunxiang, et al.
Pubblicazione: (2024)
RawGen: Learning Camera Raw Image Generation
di: Kim, Dongyoung, et al.
Pubblicazione: (2026)
di: Kim, Dongyoung, et al.
Pubblicazione: (2026)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
di: Fang, I-Sheng, et al.
Pubblicazione: (2025)
di: Fang, I-Sheng, et al.
Pubblicazione: (2025)
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
di: Fan, Weichen, et al.
Pubblicazione: (2025)
di: Fan, Weichen, et al.
Pubblicazione: (2025)
Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
di: Fu, Xiao, et al.
Pubblicazione: (2025)
di: Fu, Xiao, et al.
Pubblicazione: (2025)
Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation
di: Chen, Gordon, et al.
Pubblicazione: (2026)
di: Chen, Gordon, et al.
Pubblicazione: (2026)
Documenti analoghi
-
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
di: Yang, Shuai, et al.
Pubblicazione: (2024) -
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
di: Zhang, Yuhan, et al.
Pubblicazione: (2025) -
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
di: Wu, Tong, et al.
Pubblicazione: (2024) -
Video World Models with Long-term Spatial Memory
di: Wu, Tong, et al.
Pubblicazione: (2025) -
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
di: Wu, Tong, et al.
Pubblicazione: (2024)