CPA: Camera-pose-awareness Diffusion Transformer for Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yuelei, Zhang, Jian, Jiang, Pengtao, Zhang, Hao, Chen, Jinwei, Li, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human
von: Ying, Qijun, et al.
Veröffentlicht: (2025)
von: Ying, Qijun, et al.
Veröffentlicht: (2025)
MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
GenCompositor: Generative Video Compositing with Diffusion Transformer
von: Yang, Shuzhou, et al.
Veröffentlicht: (2025)
von: Yang, Shuzhou, et al.
Veröffentlicht: (2025)
ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model
von: Qi, Jinwei, et al.
Veröffentlicht: (2025)
von: Qi, Jinwei, et al.
Veröffentlicht: (2025)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
von: He, Hao, et al.
Veröffentlicht: (2024)
von: He, Hao, et al.
Veröffentlicht: (2024)
Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers
von: Zhu, Jingyuan, et al.
Veröffentlicht: (2026)
von: Zhu, Jingyuan, et al.
Veröffentlicht: (2026)
High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity
von: Yu, Qian, et al.
Veröffentlicht: (2024)
von: Yu, Qian, et al.
Veröffentlicht: (2024)
Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning
von: Liu, Yuti, et al.
Veröffentlicht: (2024)
von: Liu, Yuti, et al.
Veröffentlicht: (2024)
Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Framework
von: Shi, Linxiao, et al.
Veröffentlicht: (2026)
von: Shi, Linxiao, et al.
Veröffentlicht: (2026)
Improving Consistency in Diffusion Models for Image Super-Resolution
von: Gu, Junhao, et al.
Veröffentlicht: (2024)
von: Gu, Junhao, et al.
Veröffentlicht: (2024)
DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation
von: He, Xiankang, et al.
Veröffentlicht: (2024)
von: He, Xiankang, et al.
Veröffentlicht: (2024)
IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation
von: Wu, Hao, et al.
Veröffentlicht: (2026)
von: Wu, Hao, et al.
Veröffentlicht: (2026)
Boosting Camera Motion Control for Video Diffusion Transformers
von: Cheong, Soon Yau, et al.
Veröffentlicht: (2024)
von: Cheong, Soon Yau, et al.
Veröffentlicht: (2024)
Diffusion-based Data Augmentation for Object Counting Problems
von: Wang, Zhen, et al.
Veröffentlicht: (2024)
von: Wang, Zhen, et al.
Veröffentlicht: (2024)
IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control
von: Liu, Lijuan, et al.
Veröffentlicht: (2025)
von: Liu, Lijuan, et al.
Veröffentlicht: (2025)
SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
von: Zhang, Guiyu, et al.
Veröffentlicht: (2026)
von: Zhang, Guiyu, et al.
Veröffentlicht: (2026)
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention
von: Xu, Dejia, et al.
Veröffentlicht: (2024)
von: Xu, Dejia, et al.
Veröffentlicht: (2024)
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
von: Zhang, Hongfei, et al.
Veröffentlicht: (2025)
von: Zhang, Hongfei, et al.
Veröffentlicht: (2025)
CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models
von: He, Hao, et al.
Veröffentlicht: (2025)
von: He, Hao, et al.
Veröffentlicht: (2025)
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
von: Ge, Wenhang, et al.
Veröffentlicht: (2026)
von: Ge, Wenhang, et al.
Veröffentlicht: (2026)
ControlSR: Taming Diffusion Models for Consistent Real-World Image Super Resolution
von: Wan, Yuhao, et al.
Veröffentlicht: (2024)
von: Wan, Yuhao, et al.
Veröffentlicht: (2024)
Latte: Latent Diffusion Transformer for Video Generation
von: Ma, Xin, et al.
Veröffentlicht: (2024)
von: Ma, Xin, et al.
Veröffentlicht: (2024)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
von: Kuang, Zhengfei, et al.
Veröffentlicht: (2024)
von: Kuang, Zhengfei, et al.
Veröffentlicht: (2024)
SDMatte: Grafting Diffusion Models for Interactive Matting
von: Huang, Longfei, et al.
Veröffentlicht: (2025)
von: Huang, Longfei, et al.
Veröffentlicht: (2025)
Tora: Trajectory-oriented Diffusion Transformer for Video Generation
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2024)
CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
von: Wang, Qian, et al.
Veröffentlicht: (2024)
von: Wang, Qian, et al.
Veröffentlicht: (2024)
Controllable and Expressive One-Shot Video Head Swapping
von: Ji, Chaonan, et al.
Veröffentlicht: (2025)
von: Ji, Chaonan, et al.
Veröffentlicht: (2025)
BLO-Inst: Bi-Level Optimization Based Alignment of YOLO and SAM for Robust Instance Segmentation
von: Zhang, Li, et al.
Veröffentlicht: (2026)
von: Zhang, Li, et al.
Veröffentlicht: (2026)
Scalable Visual State Space Model with Fractal Scanning
von: Tang, Lv, et al.
Veröffentlicht: (2024)
von: Tang, Lv, et al.
Veröffentlicht: (2024)
Frequency-aware Neural Representation for Videos
von: Zhu, Jun, et al.
Veröffentlicht: (2026)
von: Zhu, Jun, et al.
Veröffentlicht: (2026)
Improving Adversarial Energy-Based Model via Diffusion Process
von: Geng, Cong, et al.
Veröffentlicht: (2024)
von: Geng, Cong, et al.
Veröffentlicht: (2024)
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
von: Liu, Tianqi, et al.
Veröffentlicht: (2025)
von: Liu, Tianqi, et al.
Veröffentlicht: (2025)
OneTo3D: One Image to Re-editable Dynamic 3D Model and Video Generation
von: Lin, Jinwei
Veröffentlicht: (2024)
von: Lin, Jinwei
Veröffentlicht: (2024)
GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
von: Fortier-Chouinard, Frédéric, et al.
Veröffentlicht: (2025)
von: Fortier-Chouinard, Frédéric, et al.
Veröffentlicht: (2025)
CamI2V: Camera-Controlled Image-to-Video Diffusion Model
von: Zheng, Guangcong, et al.
Veröffentlicht: (2024)
von: Zheng, Guangcong, et al.
Veröffentlicht: (2024)
MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
Multi-Task Dense Prediction via Mixture of Low-Rank Experts
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
von: Yang, Yuqi, et al.
Veröffentlicht: (2024)
Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
von: Tang, Lv, et al.
Veröffentlicht: (2023)
von: Tang, Lv, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human
von: Ying, Qijun, et al.
Veröffentlicht: (2025) -
MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
von: Li, Guangyuan, et al.
Veröffentlicht: (2025) -
GenCompositor: Generative Video Compositing with Diffusion Transformer
von: Yang, Shuzhou, et al.
Veröffentlicht: (2025) -
ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model
von: Qi, Jinwei, et al.
Veröffentlicht: (2025) -
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
von: He, Hao, et al.
Veröffentlicht: (2024)