Video Motion Transfer with Diffusion Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pondaven, Alexander, Siarohin, Aliaksandr, Tulyakov, Sergey, Torr, Philip, Pizzati, Fabio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ActionParty: Multi-Subject Action Binding in Generative Video Games
von: Pondaven, Alexander, et al.
Veröffentlicht: (2026)
von: Pondaven, Alexander, et al.
Veröffentlicht: (2026)
MatchDiffusion: Training-free Generation of Match-cuts
von: Pardo, Alejandro, et al.
Veröffentlicht: (2024)
von: Pardo, Alejandro, et al.
Veröffentlicht: (2024)
Improving the Diffusability of Autoencoders
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2025)
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2025)
Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
von: Iakovleva, Ekaterina, et al.
Veröffentlicht: (2024)
von: Iakovleva, Ekaterina, et al.
Veröffentlicht: (2024)
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
von: Khalifi, Omar El, et al.
Veröffentlicht: (2026)
von: Khalifi, Omar El, et al.
Veröffentlicht: (2026)
AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
von: Girish, Sharath, et al.
Veröffentlicht: (2025)
von: Girish, Sharath, et al.
Veröffentlicht: (2025)
Latent Guard: a Safety Framework for Text-to-image Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
Hierarchical Patch Diffusion Models for High-Resolution Video Generation
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2024)
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2024)
Promptable Game Models: Text-Guided Game Simulation via Masked Diffusion Models
von: Menapace, Willi, et al.
Veröffentlicht: (2023)
von: Menapace, Willi, et al.
Veröffentlicht: (2023)
Learning to Generate Rigid Body Interactions with Video Diffusion Models
von: Romero, David, et al.
Veröffentlicht: (2025)
von: Romero, David, et al.
Veröffentlicht: (2025)
On Pretraining Data Diversity for Self-Supervised Learning
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
von: Menapace, Willi, et al.
Veröffentlicht: (2024)
von: Menapace, Willi, et al.
Veröffentlicht: (2024)
EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
von: Li, Runjia, et al.
Veröffentlicht: (2025)
von: Li, Runjia, et al.
Veröffentlicht: (2025)
Dynamic Concepts Personalization from Single Videos
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
Spectral Motion Alignment for Video Motion Transfer using Diffusion Models
von: Park, Geon Yeong, et al.
Veröffentlicht: (2024)
von: Park, Geon Yeong, et al.
Veröffentlicht: (2024)
OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
AlphaFlow: Understanding and Improving MeanFlow Models
von: Zhang, Huijie, et al.
Veröffentlicht: (2025)
von: Zhang, Huijie, et al.
Veröffentlicht: (2025)
Improving Progressive Generation with Decomposable Flow Matching
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2025)
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2025)
Pixel-Aligned Multi-View Generation with Depth Guided Decoder
von: Tang, Zhenggang, et al.
Veröffentlicht: (2024)
von: Tang, Zhenggang, et al.
Veröffentlicht: (2024)
VIMI: Grounding Video Generation through Multi-modal Instruction
von: Fang, Yuwei, et al.
Veröffentlicht: (2024)
von: Fang, Yuwei, et al.
Veröffentlicht: (2024)
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
von: Yuan, Jianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Jianhao, et al.
Veröffentlicht: (2025)
AsCAN: Asymmetric Convolution-Attention Networks for Efficient Recognition and Generation
von: Kag, Anil, et al.
Veröffentlicht: (2024)
von: Kag, Anil, et al.
Veröffentlicht: (2024)
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2024)
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2024)
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
von: Jun, Youngjun, et al.
Veröffentlicht: (2026)
von: Jun, Youngjun, et al.
Veröffentlicht: (2026)
RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025)
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025)
PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization
von: Zhang, Yangsong, et al.
Veröffentlicht: (2026)
von: Zhang, Yangsong, et al.
Veröffentlicht: (2026)
GTR: Improving Large 3D Reconstruction Models through Geometry and Texture Refinement
von: Zhuang, Peiye, et al.
Veröffentlicht: (2024)
von: Zhuang, Peiye, et al.
Veröffentlicht: (2024)
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
von: Wu, Yen-Siang, et al.
Veröffentlicht: (2025)
von: Wu, Yen-Siang, et al.
Veröffentlicht: (2025)
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
Diffusion Priors for Dynamic View Synthesis from Monocular Videos
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
Towards Precise Scaling Laws for Video Diffusion Transformers
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
Motion meets Attention: Video Motion Prompts
von: Chen, Qixiang, et al.
Veröffentlicht: (2024)
von: Chen, Qixiang, et al.
Veröffentlicht: (2024)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
von: Burgess, James, et al.
Veröffentlicht: (2026)
von: Burgess, James, et al.
Veröffentlicht: (2026)
The Ingredients for Robotic Diffusion Transformers
von: Dasari, Sudeep, et al.
Veröffentlicht: (2024)
von: Dasari, Sudeep, et al.
Veröffentlicht: (2024)
DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
von: Wu, Ziyi, et al.
Veröffentlicht: (2025)
von: Wu, Ziyi, et al.
Veröffentlicht: (2025)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
von: Chen, Weifeng, et al.
Veröffentlicht: (2023)
von: Chen, Weifeng, et al.
Veröffentlicht: (2023)
GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
von: Song, Zhiye, et al.
Veröffentlicht: (2025)
von: Song, Zhiye, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ActionParty: Multi-Subject Action Binding in Generative Video Games
von: Pondaven, Alexander, et al.
Veröffentlicht: (2026) -
MatchDiffusion: Training-free Generation of Match-cuts
von: Pardo, Alejandro, et al.
Veröffentlicht: (2024) -
Improving the Diffusability of Autoencoders
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2025) -
Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
von: Iakovleva, Ekaterina, et al.
Veröffentlicht: (2024) -
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
von: Khalifi, Omar El, et al.
Veröffentlicht: (2026)