End-to-end Training for Text-to-Image Synthesis using Dual-Text Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmed, Yeruru Asrar, Mittal, Anurag |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncovering the Handwritten Text in the Margins: End-to-end Handwritten Text Detection and Recognition
by: Cheng, Liang, et al.
Published: (2023)
by: Cheng, Liang, et al.
Published: (2023)
SHaDe: Compact and Consistent Dynamic 3D Reconstruction via Tri-Plane Deformation and Latent Diffusion
by: Alruwayqi, Asrar
Published: (2025)
by: Alruwayqi, Asrar
Published: (2025)
RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
by: Jian, Siyong, et al.
Published: (2026)
by: Jian, Siyong, et al.
Published: (2026)
DECOR:Decomposition and Projection of Text Embeddings for Text-to-Image Customization
by: Jang, Geonhui, et al.
Published: (2024)
by: Jang, Geonhui, et al.
Published: (2024)
DiverseDream: Diverse Text-to-3D Synthesis with Augmented Text Embedding
by: Tran, Uy Dieu, et al.
Published: (2023)
by: Tran, Uy Dieu, et al.
Published: (2023)
Geometric Disentanglement of Text Embeddings for Subject-Consistent Text-to-Image Generation using A Single Prompt
by: Li, Shangxun, et al.
Published: (2025)
by: Li, Shangxun, et al.
Published: (2025)
Zero-Painter: Training-Free Layout Control for Text-to-Image Synthesis
by: Ohanyan, Marianna, et al.
Published: (2024)
by: Ohanyan, Marianna, et al.
Published: (2024)
Training-free Color-Style Disentanglement for Constrained Text-to-Image Synthesis
by: Agarwal, Aishwarya, et al.
Published: (2024)
by: Agarwal, Aishwarya, et al.
Published: (2024)
SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
by: Luo, Dongliang, et al.
Published: (2025)
by: Luo, Dongliang, et al.
Published: (2025)
CoRe: Context-Regularized Text Embedding Learning for Text-to-Image Personalization
by: Wu, Feize, et al.
Published: (2024)
by: Wu, Feize, et al.
Published: (2024)
CLUE: Controllable Latent space of Unprompted Embeddings for Diversity Management in Text-to-Image Synthesis
by: Park, Keunwoo, et al.
Published: (2025)
by: Park, Keunwoo, et al.
Published: (2025)
CountDiffusion: Text-to-Image Synthesis with Training-Free Counting-Guidance Diffusion
by: Li, Yanyu, et al.
Published: (2025)
by: Li, Yanyu, et al.
Published: (2025)
Uncovering the Text Embedding in Text-to-Image Diffusion Models
by: Yu, Hu, et al.
Published: (2024)
by: Yu, Hu, et al.
Published: (2024)
DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
End2end-ALARA: Approaching the ALARA Law in CT Imaging with End-to-end Learning
by: Tao, Xi, et al.
Published: (2025)
by: Tao, Xi, et al.
Published: (2025)
Spherical Dense Text-to-Image Synthesis
by: Winter, Timon, et al.
Published: (2025)
by: Winter, Timon, et al.
Published: (2025)
Fine-grained Text to Image Synthesis
by: Ouyang, Xu, et al.
Published: (2024)
by: Ouyang, Xu, et al.
Published: (2024)
PEO: Training-Free Aesthetic Quality Enhancement in Pre-Trained Text-to-Image Diffusion Models with Prompt Embedding Optimization
by: Margaryan, Hovhannes, et al.
Published: (2025)
by: Margaryan, Hovhannes, et al.
Published: (2025)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
by: Han, Woojung, et al.
Published: (2025)
by: Han, Woojung, et al.
Published: (2025)
PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
by: Chen, Junsong, et al.
Published: (2023)
by: Chen, Junsong, et al.
Published: (2023)
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
by: Gunawan, Agus, et al.
Published: (2025)
by: Gunawan, Agus, et al.
Published: (2025)
DeltaEdit: Exploring Text-free Training for Text-Driven Image Manipulation
by: Lyu, Yueming, et al.
Published: (2023)
by: Lyu, Yueming, et al.
Published: (2023)
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
by: Zarei, Arman, et al.
Published: (2024)
by: Zarei, Arman, et al.
Published: (2024)
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
by: Koo, Jaywon, et al.
Published: (2025)
by: Koo, Jaywon, et al.
Published: (2025)
BLENDER: Blended Text Embeddings and Diffusion Residuals for Intra-Class Image Synthesis in Deep Metric Learning
by: Kolf, Jan Niklas, et al.
Published: (2026)
by: Kolf, Jan Niklas, et al.
Published: (2026)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Text-to-Image Synthesis: A Decade Survey
by: Zhang, Nonghai, et al.
Published: (2024)
by: Zhang, Nonghai, et al.
Published: (2024)
All-in-One Conditioning for Text-to-Image Synthesis
by: Jayasekara, Hirunima, et al.
Published: (2026)
by: Jayasekara, Hirunima, et al.
Published: (2026)
TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
by: Zhai, Yukun, et al.
Published: (2023)
by: Zhai, Yukun, et al.
Published: (2023)
Addressing Text Embedding Leakage in Diffusion-based Image Editing
by: Mun, Sunung, et al.
Published: (2024)
by: Mun, Sunung, et al.
Published: (2024)
The Unreasonable Effectiveness of Text Embedding Interpolation for Continuous Image Steering
by: Ekin, Yigit, et al.
Published: (2026)
by: Ekin, Yigit, et al.
Published: (2026)
Latent Diffusion for Medical Image Segmentation: End to end learning for fast sampling and accuracy
by: Zaman, Fahim Ahmed, et al.
Published: (2024)
by: Zaman, Fahim Ahmed, et al.
Published: (2024)
DragText: Rethinking Text Embedding in Point-based Image Editing
by: Choi, Gayoon, et al.
Published: (2024)
by: Choi, Gayoon, et al.
Published: (2024)
Navigating Text-to-Image Generative Bias across Indic Languages
by: Mittal, Surbhi, et al.
Published: (2024)
by: Mittal, Surbhi, et al.
Published: (2024)
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
by: Su, Tongtong, et al.
Published: (2025)
by: Su, Tongtong, et al.
Published: (2025)
PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards
by: Le, Minh-Quan, et al.
Published: (2026)
by: Le, Minh-Quan, et al.
Published: (2026)
Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
by: Hu, Taihang, et al.
Published: (2024)
by: Hu, Taihang, et al.
Published: (2024)
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
by: Wang, Jiamian, et al.
Published: (2024)
by: Wang, Jiamian, et al.
Published: (2024)
DanceText: A Training-Free Layered Framework for Controllable Multilingual Text Transformation in Images
by: Yu, Zhenyu, et al.
Published: (2025)
by: Yu, Zhenyu, et al.
Published: (2025)
Translation of Text Embedding via Delta Vector to Suppress Strongly Entangled Content in Text-to-Image Diffusion Models
by: Koh, Eunseo, et al.
Published: (2025)
by: Koh, Eunseo, et al.
Published: (2025)
Similar Items
-
Uncovering the Handwritten Text in the Margins: End-to-end Handwritten Text Detection and Recognition
by: Cheng, Liang, et al.
Published: (2023) -
SHaDe: Compact and Consistent Dynamic 3D Reconstruction via Tri-Plane Deformation and Latent Diffusion
by: Alruwayqi, Asrar
Published: (2025) -
RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
by: Jian, Siyong, et al.
Published: (2026) -
DECOR:Decomposition and Projection of Text Embeddings for Text-to-Image Customization
by: Jang, Geonhui, et al.
Published: (2024) -
DiverseDream: Diverse Text-to-3D Synthesis with Augmented Text Embedding
by: Tran, Uy Dieu, et al.
Published: (2023)