TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xingrui, Li, Xin, Hu, Yaosi, Zhu, Hanxin, Hou, Chen, Lan, Cuiling, Chen, Zhibo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GSemSplat: Generalizable Semantic 3D Gaussian Splatting from Uncalibrated Image Pairs
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
von: Hu, Yaosi, et al.
Veröffentlicht: (2023)
von: Hu, Yaosi, et al.
Veröffentlicht: (2023)
Diffusion Models for Image Restoration and Enhancement: A Comprehensive Survey
von: Li, Xin, et al.
Veröffentlicht: (2023)
von: Li, Xin, et al.
Veröffentlicht: (2023)
UCIP: A Universal Framework for Compressed Image Super-Resolution using Dynamic Prompt
von: Li, Xin, et al.
Veröffentlicht: (2024)
von: Li, Xin, et al.
Veröffentlicht: (2024)
CoNo: Consistency Noise Injection for Tuning-free Long Video Diffusion
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion
von: Zhu, Hanxin, et al.
Veröffentlicht: (2026)
von: Zhu, Hanxin, et al.
Veröffentlicht: (2026)
GaussianSR: 3D Gaussian Super-Resolution with 2D Diffusion Priors
von: Yu, Xiqian, et al.
Veröffentlicht: (2024)
von: Yu, Xiqian, et al.
Veröffentlicht: (2024)
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
von: Wang, Cong, et al.
Veröffentlicht: (2026)
von: Wang, Cong, et al.
Veröffentlicht: (2026)
Compositional 3D-aware Video Generation with LLM Director
von: Zhu, Hanxin, et al.
Veröffentlicht: (2024)
von: Zhu, Hanxin, et al.
Veröffentlicht: (2024)
CMC: Few-shot Novel View Synthesis via Cross-view Multiplane Consistency
von: Zhu, Hanxin, et al.
Veröffentlicht: (2024)
von: Zhu, Hanxin, et al.
Veröffentlicht: (2024)
AR4D: Autoregressive 4D Generation from Monocular Videos
von: Zhu, Hanxin, et al.
Veröffentlicht: (2025)
von: Zhu, Hanxin, et al.
Veröffentlicht: (2025)
Training-free Camera Control for Video Generation
von: Hou, Chen, et al.
Veröffentlicht: (2024)
von: Hou, Chen, et al.
Veröffentlicht: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
Light Field Compression Based on Implicit Neural Representation
von: Wang, Henan, et al.
Veröffentlicht: (2024)
von: Wang, Henan, et al.
Veröffentlicht: (2024)
4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
Is Vanilla MLP in Neural Radiance Field Enough for Few-shot View Synthesis?
von: Zhu, Hanxin, et al.
Veröffentlicht: (2024)
von: Zhu, Hanxin, et al.
Veröffentlicht: (2024)
High-Fidelity Diffusion-based Image Editing
von: Hou, Chen, et al.
Veröffentlicht: (2023)
von: Hou, Chen, et al.
Veröffentlicht: (2023)
P-4DGS: Predictive 4D Gaussian Splatting with 90$\times$ Compression
von: Wang, Henan, et al.
Veröffentlicht: (2025)
von: Wang, Henan, et al.
Veröffentlicht: (2025)
Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
von: Yang, Shiyuan, et al.
Veröffentlicht: (2024)
von: Yang, Shiyuan, et al.
Veröffentlicht: (2024)
ObjectMover: Generative Object Movement with Video Prior
von: Yu, Xin, et al.
Veröffentlicht: (2025)
von: Yu, Xin, et al.
Veröffentlicht: (2025)
LiftVSR: Lifting Image Diffusion to Video Super-Resolution via Hybrid Temporal Modeling with Only 4$\times$RTX 4090s
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
von: Wang, Xijun, et al.
Veröffentlicht: (2025)
HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
von: Li, Yayuan, et al.
Veröffentlicht: (2024)
von: Li, Yayuan, et al.
Veröffentlicht: (2024)
SeD: Semantic-Aware Discriminator for Image Super-Resolution
von: Li, Bingchen, et al.
Veröffentlicht: (2024)
von: Li, Bingchen, et al.
Veröffentlicht: (2024)
Slot-VLM: SlowFast Slots for Video-Language Modeling
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
von: Zhu, Zixin, et al.
Veröffentlicht: (2024)
von: Zhu, Zixin, et al.
Veröffentlicht: (2024)
NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation
von: Feng, X., et al.
Veröffentlicht: (2025)
von: Feng, X., et al.
Veröffentlicht: (2025)
TextOCVP: Object-Centric Video Prediction with Language Guidance
von: Villar-Corrales, Angel, et al.
Veröffentlicht: (2025)
von: Villar-Corrales, Angel, et al.
Veröffentlicht: (2025)
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
von: Oh, Yoonjin, et al.
Veröffentlicht: (2025)
von: Oh, Yoonjin, et al.
Veröffentlicht: (2025)
MoE-DiffIR: Task-customized Diffusion Priors for Universal Compressed Image Restoration
von: Ren, Yulin, et al.
Veröffentlicht: (2024)
von: Ren, Yulin, et al.
Veröffentlicht: (2024)
Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image Generation
von: Wang, Junyan, et al.
Veröffentlicht: (2024)
von: Wang, Junyan, et al.
Veröffentlicht: (2024)
Object-Centric Diffusion for Efficient Video Editing
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
VideoTetris: Towards Compositional Text-to-Video Generation
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents
von: Li, Bingchen, et al.
Veröffentlicht: (2024)
von: Li, Bingchen, et al.
Veröffentlicht: (2024)
Text-driven Human Motion Generation with Motion Masked Diffusion Model
von: Chen, Xingyu
Veröffentlicht: (2024)
von: Chen, Xingyu
Veröffentlicht: (2024)
TrajectoryMover: Generative Movement of Object Trajectories in Videos
von: Chhatre, Kiran, et al.
Veröffentlicht: (2026)
von: Chhatre, Kiran, et al.
Veröffentlicht: (2026)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
von: Rahman, Tanzila, et al.
Veröffentlicht: (2024)
von: Rahman, Tanzila, et al.
Veröffentlicht: (2024)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
von: Dou, Weijia, et al.
Veröffentlicht: (2026)
von: Dou, Weijia, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
GSemSplat: Generalizable Semantic 3D Gaussian Splatting from Uncalibrated Image Pairs
von: Wang, Xingrui, et al.
Veröffentlicht: (2024) -
LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
von: Hu, Yaosi, et al.
Veröffentlicht: (2023) -
Diffusion Models for Image Restoration and Enhancement: A Comprehensive Survey
von: Li, Xin, et al.
Veröffentlicht: (2023) -
UCIP: A Universal Framework for Compressed Image Super-Resolution using Dynamic Prompt
von: Li, Xin, et al.
Veröffentlicht: (2024) -
CoNo: Consistency Noise Injection for Tuning-free Long Video Diffusion
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)