Cocktail: Mixing Multi-Modality Controls for Text-Conditional Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Minghui, Zheng, Jianbin, Liu, Daqing, Zheng, Chuanxia, Wang, Chaoyue, Tao, Dacheng, Cham, Tat-Jen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantix: An Energy Guided Sampler for Semantic Style Transfer
von: He, Huiang, et al.
Veröffentlicht: (2025)
von: He, Huiang, et al.
Veröffentlicht: (2025)
Trajectory Consistency Distillation: Improved Latent Consistency Distillation by Semi-Linear Consistency Function with Trajectory Mapping
von: Zheng, Jianbin, et al.
Veröffentlicht: (2024)
von: Zheng, Jianbin, et al.
Veröffentlicht: (2024)
SPATIALALIGN: Aligning Dynamic Spatial Relationships in Video Generation
von: Liu, Fengming, et al.
Veröffentlicht: (2026)
von: Liu, Fengming, et al.
Veröffentlicht: (2026)
One-shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing
von: Ji, Yuzhu, et al.
Veröffentlicht: (2024)
von: Ji, Yuzhu, et al.
Veröffentlicht: (2024)
PanoDiffusion: 360-degree Panorama Outpainting via Diffusion
von: Wu, Tianhao, et al.
Veröffentlicht: (2023)
von: Wu, Tianhao, et al.
Veröffentlicht: (2023)
ClusteringSDF: Self-Organized Neural Implicit Surfaces for 3D Decomposition
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
von: Wu, Tianhao, et al.
Veröffentlicht: (2025)
von: Wu, Tianhao, et al.
Veröffentlicht: (2025)
MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images
von: Chen, Yuedong, et al.
Veröffentlicht: (2024)
von: Chen, Yuedong, et al.
Veröffentlicht: (2024)
Explicit Correspondence Matching for Generalizable Neural Radiance Fields
von: Chen, Yuedong, et al.
Veröffentlicht: (2023)
von: Chen, Yuedong, et al.
Veröffentlicht: (2023)
MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views
von: Chen, Yuedong, et al.
Veröffentlicht: (2024)
von: Chen, Yuedong, et al.
Veröffentlicht: (2024)
Global Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation
von: Hu, Minghui, et al.
Veröffentlicht: (2021)
von: Hu, Minghui, et al.
Veröffentlicht: (2021)
ESceme: Vision-and-Language Navigation with Episodic Scene Memory
von: Zheng, Qi, et al.
Veröffentlicht: (2023)
von: Zheng, Qi, et al.
Veröffentlicht: (2023)
HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting
von: Lu, Wenquan, et al.
Veröffentlicht: (2023)
von: Lu, Wenquan, et al.
Veröffentlicht: (2023)
3iGS: Factorised Tensorial Illumination for 3D Gaussian Splatting
von: Tang, Zhe Jun, et al.
Veröffentlicht: (2024)
von: Tang, Zhe Jun, et al.
Veröffentlicht: (2024)
Visual Superordinate Abstraction for Robust Concept Learning
von: Zheng, Qi, et al.
Veröffentlicht: (2022)
von: Zheng, Qi, et al.
Veröffentlicht: (2022)
Towards Modality-agnostic Label-efficient Segmentation with Entropy-Regularized Distribution Alignment
von: Tang, Liyao, et al.
Veröffentlicht: (2024)
von: Tang, Liyao, et al.
Veröffentlicht: (2024)
Instilling Multi-round Thinking to Text-guided Image Generation
von: Zeng, Lidong, et al.
Veröffentlicht: (2024)
von: Zeng, Lidong, et al.
Veröffentlicht: (2024)
Free3D: Consistent Novel View Synthesis without 3D Representation
von: Zheng, Chuanxia, et al.
Veröffentlicht: (2023)
von: Zheng, Chuanxia, et al.
Veröffentlicht: (2023)
GeoConv: Geodesic Guided Convolution for Facial Action Unit Recognition
von: Chen, Yuedong, et al.
Veröffentlicht: (2020)
von: Chen, Yuedong, et al.
Veröffentlicht: (2020)
Diffusion Cocktail: Mixing Domain-Specific Diffusion Models for Diversified Image Generations
von: Liu, Haoming, et al.
Veröffentlicht: (2023)
von: Liu, Haoming, et al.
Veröffentlicht: (2023)
Scaling Down Text Encoders of Text-to-Image Diffusion Models
von: Wang, Lifu, et al.
Veröffentlicht: (2025)
von: Wang, Lifu, et al.
Veröffentlicht: (2025)
Connecting Consistency Distillation to Score Distillation for Text-to-3D Generation
von: Li, Zongrui, et al.
Veröffentlicht: (2024)
von: Li, Zongrui, et al.
Veröffentlicht: (2024)
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
von: Xuan, Wenjie, et al.
Veröffentlicht: (2024)
von: Xuan, Wenjie, et al.
Veröffentlicht: (2024)
D-LORD for Motion Stylization
von: Gupta, Meenakshi, et al.
Veröffentlicht: (2024)
von: Gupta, Meenakshi, et al.
Veröffentlicht: (2024)
SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation
von: Ma, Yingzi, et al.
Veröffentlicht: (2026)
von: Ma, Yingzi, et al.
Veröffentlicht: (2026)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
von: Chen, Yiyang, et al.
Veröffentlicht: (2022)
von: Chen, Yiyang, et al.
Veröffentlicht: (2022)
Local Conditional Controlling for Text-to-Image Diffusion Models
von: Zhao, Yibo, et al.
Veröffentlicht: (2023)
von: Zhao, Yibo, et al.
Veröffentlicht: (2023)
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
von: He, Qingdong, et al.
Veröffentlicht: (2024)
von: He, Qingdong, et al.
Veröffentlicht: (2024)
Reverse Prompt: Cracking the Recipe Inside Text-to-Image Generation
von: Ren, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ren, Zhiyao, et al.
Veröffentlicht: (2025)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
von: Zheng, Zirui, et al.
Veröffentlicht: (2025)
von: Zheng, Zirui, et al.
Veröffentlicht: (2025)
Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level Dynamics
von: Li, Ruining, et al.
Veröffentlicht: (2024)
von: Li, Ruining, et al.
Veröffentlicht: (2024)
Amodal Ground Truth and Completion in the Wild
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
DragAPart: Learning a Part-Level Motion Prior for Articulated Objects
von: Li, Ruining, et al.
Veröffentlicht: (2024)
von: Li, Ruining, et al.
Veröffentlicht: (2024)
Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs
von: Smart, Brandon, et al.
Veröffentlicht: (2024)
von: Smart, Brandon, et al.
Veröffentlicht: (2024)
Rethink Sparse Signals for Pose-guided Text-to-image Generation
von: Xuan, Wenjie, et al.
Veröffentlicht: (2025)
von: Xuan, Wenjie, et al.
Veröffentlicht: (2025)
MultiDiffSense: Diffusion-Based Multi-Modal Visuo-Tactile Image Generation Conditioned on Object Shape and Contact Pose
von: Bhouri, Sirine, et al.
Veröffentlicht: (2026)
von: Bhouri, Sirine, et al.
Veröffentlicht: (2026)
Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective
von: Wang, Weijie, et al.
Veröffentlicht: (2026)
von: Wang, Weijie, et al.
Veröffentlicht: (2026)
MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling
von: Yang, Jian, et al.
Veröffentlicht: (2024)
von: Yang, Jian, et al.
Veröffentlicht: (2024)
CogMorph: Cognitive Morphing Attacks for Text-to-Image Models
von: Jing, Zonglei, et al.
Veröffentlicht: (2025)
von: Jing, Zonglei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Semantix: An Energy Guided Sampler for Semantic Style Transfer
von: He, Huiang, et al.
Veröffentlicht: (2025) -
Trajectory Consistency Distillation: Improved Latent Consistency Distillation by Semi-Linear Consistency Function with Trajectory Mapping
von: Zheng, Jianbin, et al.
Veröffentlicht: (2024) -
SPATIALALIGN: Aligning Dynamic Spatial Relationships in Video Generation
von: Liu, Fengming, et al.
Veröffentlicht: (2026) -
One-shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing
von: Ji, Yuzhu, et al.
Veröffentlicht: (2024) -
PanoDiffusion: 360-degree Panorama Outpainting via Diffusion
von: Wu, Tianhao, et al.
Veröffentlicht: (2023)