Contextualized Diffusion Models for Text-Guided Image and Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Ling, Zhang, Zhilong, Yu, Zhaochen, Liu, Jingwei, Xu, Minkai, Ermon, Stefano, Cui, Bin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
Structure-Guided Adversarial Training of Diffusion Models
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models
von: Zhang, Xinchen, et al.
Veröffentlicht: (2024)
von: Zhang, Xinchen, et al.
Veröffentlicht: (2024)
Consistency Flow Matching: Defining Straight Flows with Velocity Consistency
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
Improving Diffusion-Based Image Synthesis with Context Prediction
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
Geometric Trajectory Diffusion Models
von: Han, Jiaqi, et al.
Veröffentlicht: (2024)
von: Han, Jiaqi, et al.
Veröffentlicht: (2024)
Divergence Minimization Preference Optimization for Diffusion Model Alignment
von: Li, Binxu, et al.
Veröffentlicht: (2025)
von: Li, Binxu, et al.
Veröffentlicht: (2025)
On the Scalability of Diffusion-based Text-to-Image Generation
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
Paired Image Generation with Diffusion-Guided Diffusion Models
von: Zhang, Haoxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Haoxuan, et al.
Veröffentlicht: (2025)
VideoTetris: Towards Compositional Text-to-Video Generation
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
SVGDreamer: Text Guided SVG Generation with Diffusion Model
von: Xing, Ximing, et al.
Veröffentlicht: (2023)
von: Xing, Ximing, et al.
Veröffentlicht: (2023)
Diffusion Models: A Comprehensive Survey of Methods and Applications
von: Yang, Ling, et al.
Veröffentlicht: (2022)
von: Yang, Ling, et al.
Veröffentlicht: (2022)
HARIVO: Harnessing Text-to-Image Models for Video Generation
von: Kwon, Mingi, et al.
Veröffentlicht: (2024)
von: Kwon, Mingi, et al.
Veröffentlicht: (2024)
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
von: Yi, Jingwei, et al.
Veröffentlicht: (2025)
von: Yi, Jingwei, et al.
Veröffentlicht: (2025)
Repairing Catastrophic-Neglect in Text-to-Image Diffusion Models via Attention-Guided Feature Enhancement
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Chang, Zhiyuan, et al.
Veröffentlicht: (2024)
LDEdit: Towards Generalized Text Guided Image Manipulation via Latent Diffusion Models
von: Chandramouli, Paramanand, et al.
Veröffentlicht: (2022)
von: Chandramouli, Paramanand, et al.
Veröffentlicht: (2022)
DiffusionSat: A Generative Foundation Model for Satellite Imagery
von: Khanna, Samar, et al.
Veröffentlicht: (2023)
von: Khanna, Samar, et al.
Veröffentlicht: (2023)
Video Text Preservation with Synthetic Text-Rich Videos
von: Liu, Ziyang, et al.
Veröffentlicht: (2025)
von: Liu, Ziyang, et al.
Veröffentlicht: (2025)
Trans4D: Realistic Geometry-Aware Transition for Compositional Text-to-4D Synthesis
von: Zeng, Bohan, et al.
Veröffentlicht: (2024)
von: Zeng, Bohan, et al.
Veröffentlicht: (2024)
Uncovering the Text Embedding in Text-to-Image Diffusion Models
von: Yu, Hu, et al.
Veröffentlicht: (2024)
von: Yu, Hu, et al.
Veröffentlicht: (2024)
GVD: Guiding Video Diffusion Model for Scalable Video Distillation
von: Li, Kunyang, et al.
Veröffentlicht: (2025)
von: Li, Kunyang, et al.
Veröffentlicht: (2025)
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
Preacher: Paper-to-Video Agentic System
von: Liu, Jingwei, et al.
Veröffentlicht: (2025)
von: Liu, Jingwei, et al.
Veröffentlicht: (2025)
Progressive Image Restoration via Text-Conditioned Video Generation
von: Kang, Peng, et al.
Veröffentlicht: (2025)
von: Kang, Peng, et al.
Veröffentlicht: (2025)
SurGen: Text-Guided Diffusion Model for Surgical Video Generation
von: Cho, Joseph, et al.
Veröffentlicht: (2024)
von: Cho, Joseph, et al.
Veröffentlicht: (2024)
DistillKac: Few-Step Image Generation via Damped Wave Equations
von: Han, Weiqiao, et al.
Veröffentlicht: (2025)
von: Han, Weiqiao, et al.
Veröffentlicht: (2025)
ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
von: Mai, Ziyang, et al.
Veröffentlicht: (2025)
von: Mai, Ziyang, et al.
Veröffentlicht: (2025)
DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models
von: Xing, Ximing, et al.
Veröffentlicht: (2023)
von: Xing, Ximing, et al.
Veröffentlicht: (2023)
G2D2: Gradient-Guided Discrete Diffusion for Inverse Problem Solving
von: Murata, Naoki, et al.
Veröffentlicht: (2024)
von: Murata, Naoki, et al.
Veröffentlicht: (2024)
EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation
von: Han, Shuhao, et al.
Veröffentlicht: (2024)
von: Han, Shuhao, et al.
Veröffentlicht: (2024)
Instant Preference Alignment for Text-to-Image Diffusion Models
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
von: Yang, Lehan, et al.
Veröffentlicht: (2025)
von: Yang, Lehan, et al.
Veröffentlicht: (2025)
BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
von: Shi, Fengyuan, et al.
Veröffentlicht: (2023)
von: Shi, Fengyuan, et al.
Veröffentlicht: (2023)
Grounding Text-to-Image Diffusion Models for Controlled High-Quality Image Generation
von: Süleyman, Ahmad, et al.
Veröffentlicht: (2025)
von: Süleyman, Ahmad, et al.
Veröffentlicht: (2025)
Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models
von: Wu, Xiaoshi, et al.
Veröffentlicht: (2024)
von: Wu, Xiaoshi, et al.
Veröffentlicht: (2024)
Pre-Trained Video Generative Models as World Simulators
von: He, Haoran, et al.
Veröffentlicht: (2025)
von: He, Haoran, et al.
Veröffentlicht: (2025)
Learning Plug-and-play Memory for Guiding Video Diffusion Models
von: Song, Selena, et al.
Veröffentlicht: (2025)
von: Song, Selena, et al.
Veröffentlicht: (2025)
VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPs
von: Yang, Ling, et al.
Veröffentlicht: (2023)
von: Yang, Ling, et al.
Veröffentlicht: (2023)
SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
von: Yang, Ling, et al.
Veröffentlicht: (2024) -
Structure-Guided Adversarial Training of Diffusion Models
von: Yang, Ling, et al.
Veröffentlicht: (2024) -
RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models
von: Zhang, Xinchen, et al.
Veröffentlicht: (2024) -
Consistency Flow Matching: Defining Straight Flows with Velocity Consistency
von: Yang, Ling, et al.
Veröffentlicht: (2024) -
Improving Diffusion-Based Image Synthesis with Context Prediction
von: Yang, Ling, et al.
Veröffentlicht: (2024)