Seek for Incantations: Towards Accurate Text-to-Image Diffusion Synthesis through Prompt Engineering
Fuente:
arXiv
Guardado en:
| Autores principales: | Yu, Chang, Peng, Junran, Zhu, Xiangyu, Zhang, Zhaoxiang, Tian, Qi, Lei, Zhen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes
por: Liu, Yang, et al.
Publicado: (2024)
por: Liu, Yang, et al.
Publicado: (2024)
EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models
por: Zhang, Yixuan, et al.
Publicado: (2025)
por: Zhang, Yixuan, et al.
Publicado: (2025)
GSPlane: Concise and Accurate Planar Reconstruction via Structured Representation
por: Gan, Ruitong, et al.
Publicado: (2025)
por: Gan, Ruitong, et al.
Publicado: (2025)
VGGT-X: When VGGT Meets Dense Novel View Synthesis
por: Liu, Yang, et al.
Publicado: (2025)
por: Liu, Yang, et al.
Publicado: (2025)
DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer
por: Ma, Zhiyuan, et al.
Publicado: (2024)
por: Ma, Zhiyuan, et al.
Publicado: (2024)
Top-Down Guidance for Learning Object-Centric Representations
por: Zou, Junhong, et al.
Publicado: (2024)
por: Zou, Junhong, et al.
Publicado: (2024)
Revisiting Marr in Face: The Building of 2D--2.5D--3D Representations in Deep Neural Networks
por: Zhu, Xiangyu, et al.
Publicado: (2024)
por: Zhu, Xiangyu, et al.
Publicado: (2024)
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
por: Zhu, Shangwen, et al.
Publicado: (2026)
por: Zhu, Shangwen, et al.
Publicado: (2026)
OOD-HOI: Text-Driven 3D Whole-Body Human-Object Interactions Generation Beyond Training Domains
por: Zhang, Yixuan, et al.
Publicado: (2024)
por: Zhang, Yixuan, et al.
Publicado: (2024)
StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
por: Huang, Yiheng, et al.
Publicado: (2024)
por: Huang, Yiheng, et al.
Publicado: (2024)
Adaptive 3D Convolution for Remote Sensing Image Fusion
por: Peng, Siran, et al.
Publicado: (2026)
por: Peng, Siran, et al.
Publicado: (2026)
Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
por: Feng, Yutong, et al.
Publicado: (2023)
por: Feng, Yutong, et al.
Publicado: (2023)
Generating on Generated: An Approach Towards Self-Evolving Diffusion Models
por: Zhang, Xulu, et al.
Publicado: (2025)
por: Zhang, Xulu, et al.
Publicado: (2025)
SAGD: Boundary-Enhanced Segment Anything in 3D Gaussian via Gaussian Decomposition
por: Hu, Xu, et al.
Publicado: (2024)
por: Hu, Xu, et al.
Publicado: (2024)
Towards Realistic Hand-Object Interaction with Gravity-Field Based Diffusion Bridge
por: Xu, Miao, et al.
Publicado: (2025)
por: Xu, Miao, et al.
Publicado: (2025)
Generative Active Learning for Image Synthesis Personalization
por: Zhang, Xulu, et al.
Publicado: (2024)
por: Zhang, Xulu, et al.
Publicado: (2024)
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
por: Bai, Jinbin, et al.
Publicado: (2024)
por: Bai, Jinbin, et al.
Publicado: (2024)
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
por: Yang, Yuxue, et al.
Publicado: (2026)
por: Yang, Yuxue, et al.
Publicado: (2026)
LesionDiffusion: Towards Text-controlled General Lesion Synthesis
por: Lei, Wenhui, et al.
Publicado: (2025)
por: Lei, Wenhui, et al.
Publicado: (2025)
TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts
por: Zhuang, Jingyu, et al.
Publicado: (2024)
por: Zhuang, Jingyu, et al.
Publicado: (2024)
A Survey on Personalized Content Synthesis with Diffusion Models
por: Zhang, Xulu, et al.
Publicado: (2024)
por: Zhang, Xulu, et al.
Publicado: (2024)
ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation
por: Ma, Zhiyuan, et al.
Publicado: (2024)
por: Ma, Zhiyuan, et al.
Publicado: (2024)
Image Super-Resolution with Text Prompt Diffusion
por: Chen, Zheng, et al.
Publicado: (2023)
por: Chen, Zheng, et al.
Publicado: (2023)
Modeling Spoof Noise by De-spoofing Diffusion and its Application in Face Anti-spoofing
por: Zhang, Bin, et al.
Publicado: (2024)
por: Zhang, Bin, et al.
Publicado: (2024)
Improving Large Vision-Language Models' Understanding for Flow Field Data
por: Zhang, Xiaomei, et al.
Publicado: (2025)
por: Zhang, Xiaomei, et al.
Publicado: (2025)
FlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models
por: Zhao, Lin, et al.
Publicado: (2024)
por: Zhao, Lin, et al.
Publicado: (2024)
Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis
por: Jelaca, Aleksa, et al.
Publicado: (2025)
por: Jelaca, Aleksa, et al.
Publicado: (2025)
Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
por: Li, Yiheng, et al.
Publicado: (2025)
por: Li, Yiheng, et al.
Publicado: (2025)
CityGaussian: Real-time High-quality Large-Scale Scene Rendering with Gaussians
por: Liu, Yang, et al.
Publicado: (2024)
por: Liu, Yang, et al.
Publicado: (2024)
Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Data
por: Ma, Zhiyuan, et al.
Publicado: (2025)
por: Ma, Zhiyuan, et al.
Publicado: (2025)
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
por: Egiazarian, Vage, et al.
Publicado: (2024)
por: Egiazarian, Vage, et al.
Publicado: (2024)
Noise Diffusion for Enhancing Semantic Faithfulness in Text-to-Image Synthesis
por: Miao, Boming, et al.
Publicado: (2024)
por: Miao, Boming, et al.
Publicado: (2024)
Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
por: Wu, Xiangyu, et al.
Publicado: (2025)
por: Wu, Xiangyu, et al.
Publicado: (2025)
Debiasing Text-to-Image Diffusion Models
por: He, Ruifei, et al.
Publicado: (2024)
por: He, Ruifei, et al.
Publicado: (2024)
Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models
por: Wu, Chen, et al.
Publicado: (2024)
por: Wu, Chen, et al.
Publicado: (2024)
Towards More Accurate Diffusion Model Acceleration with A Timestep Tuner
por: Xia, Mengfei, et al.
Publicado: (2023)
por: Xia, Mengfei, et al.
Publicado: (2023)
DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
por: Peng, Siran, et al.
Publicado: (2025)
por: Peng, Siran, et al.
Publicado: (2025)
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
por: Zhang, Zhuomeng, et al.
Publicado: (2024)
por: Zhang, Zhuomeng, et al.
Publicado: (2024)
TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt
por: Wu, Xiangyu, et al.
Publicado: (2024)
por: Wu, Xiangyu, et al.
Publicado: (2024)
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
por: Zhangli, Qilong, et al.
Publicado: (2024)
por: Zhangli, Qilong, et al.
Publicado: (2024)
Ejemplares similares
-
CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes
por: Liu, Yang, et al.
Publicado: (2024) -
EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models
por: Zhang, Yixuan, et al.
Publicado: (2025) -
GSPlane: Concise and Accurate Planar Reconstruction via Structured Representation
por: Gan, Ruitong, et al.
Publicado: (2025) -
VGGT-X: When VGGT Meets Dense Novel View Synthesis
por: Liu, Yang, et al.
Publicado: (2025) -
DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer
por: Ma, Zhiyuan, et al.
Publicado: (2024)