3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, JoungBin, Jung, Jaewoo, Han, Jisang, Narihira, Takuya, Fukuda, Kazumi, Seo, Junyoung, Hong, Sunghwan, Mitsufuji, Yuki, Kim, Seungryong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
D$^2$USt3R: Enhancing 3D Reconstruction for Dynamic Scenes
von: Han, Jisang, et al.
Veröffentlicht: (2025)
von: Han, Jisang, et al.
Veröffentlicht: (2025)
Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
von: Seo, Junyoung, et al.
Veröffentlicht: (2025)
von: Seo, Junyoung, et al.
Veröffentlicht: (2025)
C3G: Learning Compact 3D Representations with 2K Gaussians
von: An, Honggyu, et al.
Veröffentlicht: (2025)
von: An, Honggyu, et al.
Veröffentlicht: (2025)
Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction
von: Kim, Mungyeom, et al.
Veröffentlicht: (2026)
von: Kim, Mungyeom, et al.
Veröffentlicht: (2026)
GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping
von: Seo, Junyoung, et al.
Veröffentlicht: (2024)
von: Seo, Junyoung, et al.
Veröffentlicht: (2024)
Cross-View Completion Models are Zero-shot Correspondence Estimators
von: An, Honggyu, et al.
Veröffentlicht: (2024)
von: An, Honggyu, et al.
Veröffentlicht: (2024)
HumanGif: Single-View Human Diffusion with Generative Prior
von: Hu, Shoukang, et al.
Veröffentlicht: (2025)
von: Hu, Shoukang, et al.
Veröffentlicht: (2025)
PF3plat: Pose-Free Feed-Forward 3D Gaussian Splatting
von: Hong, Sunghwan, et al.
Veröffentlicht: (2024)
von: Hong, Sunghwan, et al.
Veröffentlicht: (2024)
DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation
von: Hong, Susung, et al.
Veröffentlicht: (2023)
von: Hong, Susung, et al.
Veröffentlicht: (2023)
Emergent Outlier View Rejection in Visual Geometry Grounded Transformers
von: Han, Jisang, et al.
Veröffentlicht: (2025)
von: Han, Jisang, et al.
Veröffentlicht: (2025)
Relaxing Accurate Initialization Constraint for 3D Gaussian Splatting
von: Jung, Jaewoo, et al.
Veröffentlicht: (2024)
von: Jung, Jaewoo, et al.
Veröffentlicht: (2024)
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
von: Gröpl, Marcel, et al.
Veröffentlicht: (2026)
von: Gröpl, Marcel, et al.
Veröffentlicht: (2026)
Pose-dIVE: Pose-Diversified Augmentation with Diffusion Model for Person Re-Identification
von: Kim, Inès Hyeonsu, et al.
Veröffentlicht: (2024)
von: Kim, Inès Hyeonsu, et al.
Veröffentlicht: (2024)
Unifying Correspondence, Pose and NeRF for Pose-Free Novel View Synthesis from Stereo Pairs
von: Hong, Sunghwan, et al.
Veröffentlicht: (2023)
von: Hong, Sunghwan, et al.
Veröffentlicht: (2023)
WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation
von: Nam, Jisu, et al.
Veröffentlicht: (2026)
von: Nam, Jisu, et al.
Veröffentlicht: (2026)
Grounding World Simulation Models in a Real-World Metropolis
von: Seo, Junyoung, et al.
Veröffentlicht: (2026)
von: Seo, Junyoung, et al.
Veröffentlicht: (2026)
Visual Representation Alignment for Multimodal Large Language Models
von: Yoon, Heeji, et al.
Veröffentlicht: (2025)
von: Yoon, Heeji, et al.
Veröffentlicht: (2025)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
Towards Generalizable Scene Change Detection
von: Kim, Jaewoo, et al.
Veröffentlicht: (2024)
von: Kim, Jaewoo, et al.
Veröffentlicht: (2024)
APPLE: Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping
von: Kang, Jiwon, et al.
Veröffentlicht: (2026)
von: Kang, Jiwon, et al.
Veröffentlicht: (2026)
IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control
von: Liu, Lijuan, et al.
Veröffentlicht: (2025)
von: Liu, Lijuan, et al.
Veröffentlicht: (2025)
Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
von: Hong, Sunghwan, et al.
Veröffentlicht: (2024)
von: Hong, Sunghwan, et al.
Veröffentlicht: (2024)
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
von: Yuan, Yu, et al.
Veröffentlicht: (2024)
von: Yuan, Yu, et al.
Veröffentlicht: (2024)
Scene123: One Prompt to 3D Scene Generation via Video-Assisted and Consistency-Enhanced MAE
von: Yang, Yiying, et al.
Veröffentlicht: (2024)
von: Yang, Yiying, et al.
Veröffentlicht: (2024)
CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation
von: Cho, Seokju, et al.
Veröffentlicht: (2023)
von: Cho, Seokju, et al.
Veröffentlicht: (2023)
URECA: Unique Region Caption Anything
von: Lim, Sangbeom, et al.
Veröffentlicht: (2025)
von: Lim, Sangbeom, et al.
Veröffentlicht: (2025)
Domain Generalization Using Large Pretrained Models with Mixture-of-Adapters
von: Lee, Gyuseong, et al.
Veröffentlicht: (2023)
von: Lee, Gyuseong, et al.
Veröffentlicht: (2023)
Generative Photographic Control for Scene-Consistent Video Cinematic Editing
von: Sun, Huiqiang, et al.
Veröffentlicht: (2025)
von: Sun, Huiqiang, et al.
Veröffentlicht: (2025)
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
von: Nam, Jisu, et al.
Veröffentlicht: (2026)
von: Nam, Jisu, et al.
Veröffentlicht: (2026)
LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical Study
von: Yang, Dongil, et al.
Veröffentlicht: (2025)
von: Yang, Dongil, et al.
Veröffentlicht: (2025)
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels
von: Shin, Heeseong, et al.
Veröffentlicht: (2024)
von: Shin, Heeseong, et al.
Veröffentlicht: (2024)
Person-In-Situ: Scene-Consistent Human Image Insertion with Occlusion-Aware Pose Control
von: Masuda, Shun, et al.
Veröffentlicht: (2025)
von: Masuda, Shun, et al.
Veröffentlicht: (2025)
Person‐In‐Situ: Scene‐Consistent Human Image Insertion With Occlusion‐Aware Pose Control
von: Shun Masuda, et al.
Veröffentlicht: (2026)
von: Shun Masuda, et al.
Veröffentlicht: (2026)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
von: Chu, Sanghyeok, et al.
Veröffentlicht: (2025)
von: Chu, Sanghyeok, et al.
Veröffentlicht: (2025)
V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping
von: Lee, Hyunkoo, et al.
Veröffentlicht: (2025)
von: Lee, Hyunkoo, et al.
Veröffentlicht: (2025)
Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation
von: Seo, Junyoung, et al.
Veröffentlicht: (2023)
von: Seo, Junyoung, et al.
Veröffentlicht: (2023)
Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion
von: Zhou, Zhenghong, et al.
Veröffentlicht: (2026)
von: Zhou, Zhenghong, et al.
Veröffentlicht: (2026)
Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
von: Kim, Chaehyun, et al.
Veröffentlicht: (2025)
von: Kim, Chaehyun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
D$^2$USt3R: Enhancing 3D Reconstruction for Dynamic Scenes
von: Han, Jisang, et al.
Veröffentlicht: (2025) -
Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
von: Seo, Junyoung, et al.
Veröffentlicht: (2025) -
C3G: Learning Compact 3D Representations with 2K Gaussians
von: An, Honggyu, et al.
Veröffentlicht: (2025) -
Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction
von: Kim, Mungyeom, et al.
Veröffentlicht: (2026) -
GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping
von: Seo, Junyoung, et al.
Veröffentlicht: (2024)