Searching Priors Makes Text-to-Video Synthesis Better
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Haoran, Peng, Liang, Xia, Linxuan, Hu, Yuepeng, Li, Hengjia, Lu, Qinglin, He, Xiaofei, Wu, Boxi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SmoothVideo: Smooth Video Synthesis with Noise Constraints on Diffusion Models for One-shot Video Tuning
von: Peng, Liang, et al.
Veröffentlicht: (2023)
von: Peng, Liang, et al.
Veröffentlicht: (2023)
PhyRPR: Training-Free Physics-Constrained Video Generation
von: Zhao, Yibo, et al.
Veröffentlicht: (2026)
von: Zhao, Yibo, et al.
Veröffentlicht: (2026)
SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
von: Peng, Liang, et al.
Veröffentlicht: (2025)
von: Peng, Liang, et al.
Veröffentlicht: (2025)
Local Conditional Controlling for Text-to-Image Diffusion Models
von: Zhao, Yibo, et al.
Veröffentlicht: (2023)
von: Zhao, Yibo, et al.
Veröffentlicht: (2023)
Discriminator-Free Direct Preference Optimization for Video Diffusion
von: Cheng, Haoran, et al.
Veröffentlicht: (2025)
von: Cheng, Haoran, et al.
Veröffentlicht: (2025)
GCA-3D: Towards Generalized and Consistent Domain Adaptation of 3D Generators
von: Li, Hengjia, et al.
Veröffentlicht: (2024)
von: Li, Hengjia, et al.
Veröffentlicht: (2024)
LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training-Free Diffusion Models
von: Yang, Yang, et al.
Veröffentlicht: (2024)
von: Yang, Yang, et al.
Veröffentlicht: (2024)
MagicView: Multi-View Consistent Identity Customization via Priors-Guided In-Context Learning
von: Li, Hengjia, et al.
Veröffentlicht: (2025)
von: Li, Hengjia, et al.
Veröffentlicht: (2025)
MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization
von: Li, Hengjia, et al.
Veröffentlicht: (2025)
von: Li, Hengjia, et al.
Veröffentlicht: (2025)
RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
von: He, Xuming, et al.
Veröffentlicht: (2025)
von: He, Xuming, et al.
Veröffentlicht: (2025)
PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation
von: Li, Hengjia, et al.
Veröffentlicht: (2024)
von: Li, Hengjia, et al.
Veröffentlicht: (2024)
Pack and Force Your Memory: Long-form and Consistent Video Generation
von: Wu, Xiaofei, et al.
Veröffentlicht: (2025)
von: Wu, Xiaofei, et al.
Veröffentlicht: (2025)
HuPrior3R: Incorporating Human Priors for Better 3D Dynamic Reconstruction from Monocular Videos
von: Xiong, Weitao, et al.
Veröffentlicht: (2025)
von: Xiong, Weitao, et al.
Veröffentlicht: (2025)
Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
von: Yang, Yang, et al.
Veröffentlicht: (2025)
von: Yang, Yang, et al.
Veröffentlicht: (2025)
DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior
von: Huang, Junjia, et al.
Veröffentlicht: (2026)
von: Huang, Junjia, et al.
Veröffentlicht: (2026)
ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing
von: Li, Hengjia, et al.
Veröffentlicht: (2026)
von: Li, Hengjia, et al.
Veröffentlicht: (2026)
Motion-aware Memory Network for Fast Video Salient Object Detection
von: Zhao, Xing, et al.
Veröffentlicht: (2022)
von: Zhao, Xing, et al.
Veröffentlicht: (2022)
SafeText: Safe Text-to-image Models via Aligning the Text Encoder
von: Hu, Yuepeng, et al.
Veröffentlicht: (2025)
von: Hu, Yuepeng, et al.
Veröffentlicht: (2025)
Towards A Better Metric for Text-to-Video Generation
von: Wu, Jay Zhangjie, et al.
Veröffentlicht: (2024)
von: Wu, Jay Zhangjie, et al.
Veröffentlicht: (2024)
USV: Unified Sparsification for Accelerating Video Diffusion Models
von: Wu, Xinjian, et al.
Veröffentlicht: (2025)
von: Wu, Xinjian, et al.
Veröffentlicht: (2025)
Rejection Sampling IMLE: Designing Priors for Better Few-Shot Image Synthesis
von: Vashist, Chirag, et al.
Veröffentlicht: (2024)
von: Vashist, Chirag, et al.
Veröffentlicht: (2024)
Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-dataset 3D Object Detection
von: Zhang, Zhanwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zhanwei, et al.
Veröffentlicht: (2024)
AniClipart: Clipart Animation with Text-to-Video Priors
von: Wu, Ronghuan, et al.
Veröffentlicht: (2024)
von: Wu, Ronghuan, et al.
Veröffentlicht: (2024)
Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
von: Yang, Yang, et al.
Veröffentlicht: (2026)
von: Yang, Yang, et al.
Veröffentlicht: (2026)
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
von: Hu, Teng, et al.
Veröffentlicht: (2025)
von: Hu, Teng, et al.
Veröffentlicht: (2025)
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
von: Su, Tongtong, et al.
Veröffentlicht: (2025)
von: Su, Tongtong, et al.
Veröffentlicht: (2025)
OrbitNVS: Harnessing Video Diffusion Priors for Novel View Synthesis
von: Liang, Jinglin, et al.
Veröffentlicht: (2026)
von: Liang, Jinglin, et al.
Veröffentlicht: (2026)
FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
von: Wang, Yuanzhi, et al.
Veröffentlicht: (2026)
von: Wang, Yuanzhi, et al.
Veröffentlicht: (2026)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
von: Chen, Weifeng, et al.
Veröffentlicht: (2023)
von: Chen, Weifeng, et al.
Veröffentlicht: (2023)
Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
von: Hu, Taihang, et al.
Veröffentlicht: (2024)
von: Hu, Taihang, et al.
Veröffentlicht: (2024)
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
von: Lin, Yuqi, et al.
Veröffentlicht: (2025)
von: Lin, Yuqi, et al.
Veröffentlicht: (2025)
MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
von: Bai, Xiangyu, et al.
Veröffentlicht: (2025)
von: Bai, Xiangyu, et al.
Veröffentlicht: (2025)
Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration
von: Bai, Haoran, et al.
Veröffentlicht: (2025)
von: Bai, Haoran, et al.
Veröffentlicht: (2025)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
PD-APE: A Parallel Decoding Framework with Adaptive Position Encoding for 3D Visual Grounding
von: Hou, Chenshu, et al.
Veröffentlicht: (2024)
von: Hou, Chenshu, et al.
Veröffentlicht: (2024)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
von: Zi, Bojia, et al.
Veröffentlicht: (2024)
von: Zi, Bojia, et al.
Veröffentlicht: (2024)
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment
von: Jiang, Lifan, et al.
Veröffentlicht: (2025)
von: Jiang, Lifan, et al.
Veröffentlicht: (2025)
DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
von: Huang, Minbin, et al.
Veröffentlicht: (2024)
von: Huang, Minbin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SmoothVideo: Smooth Video Synthesis with Noise Constraints on Diffusion Models for One-shot Video Tuning
von: Peng, Liang, et al.
Veröffentlicht: (2023) -
PhyRPR: Training-Free Physics-Constrained Video Generation
von: Zhao, Yibo, et al.
Veröffentlicht: (2026) -
SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
von: Peng, Liang, et al.
Veröffentlicht: (2025) -
Local Conditional Controlling for Text-to-Image Diffusion Models
von: Zhao, Yibo, et al.
Veröffentlicht: (2023) -
Discriminator-Free Direct Preference Optimization for Video Diffusion
von: Cheng, Haoran, et al.
Veröffentlicht: (2025)