Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yu, Fei, Hao, Li, Xiangtai, Qin, Libo, Ji, Jiayi, Zhu, Hongyuan, Zhang, Meishan, Zhang, Min, Wei, Jianguo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Words Smile: Generating Diverse Emotional Facial Expressions from Text
by: Xu, Haidong, et al.
Published: (2024)
by: Xu, Haidong, et al.
Published: (2024)
Towards Text-Image Interleaved Retrieval
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
TIPS: Text-Image Pretraining with Spatial awareness
by: Maninis, Kevis-Kokitsi, et al.
Published: (2024)
by: Maninis, Kevis-Kokitsi, et al.
Published: (2024)
TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation
by: Wang, Alex Jinpeng, et al.
Published: (2025)
by: Wang, Alex Jinpeng, et al.
Published: (2025)
Dynamic Frequency Modulation for Controllable Text-driven Image Generation
by: Shi, Tiandong, et al.
Published: (2026)
by: Shi, Tiandong, et al.
Published: (2026)
Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Spatial-aware Symmetric Alignment for Text-guided Medical Image Segmentation
by: Liao, Linglin, et al.
Published: (2025)
by: Liao, Linglin, et al.
Published: (2025)
Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
by: Zhang, Yifei, et al.
Published: (2025)
by: Zhang, Yifei, et al.
Published: (2025)
Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
by: Yi, Xunpeng, et al.
Published: (2024)
by: Yi, Xunpeng, et al.
Published: (2024)
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
by: Bai, Jinbin, et al.
Published: (2024)
by: Bai, Jinbin, et al.
Published: (2024)
Adaptively Clustering Neighbor Elements for Image-Text Generation
by: Wang, Zihua, et al.
Published: (2023)
by: Wang, Zihua, et al.
Published: (2023)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
by: Fei, Hao, et al.
Published: (2023)
by: Fei, Hao, et al.
Published: (2023)
ACCORD: Alleviating Concept Coupling through Dependence Regularization for Text-to-Image Diffusion Personalization
by: Liu, Shizhan, et al.
Published: (2025)
by: Liu, Shizhan, et al.
Published: (2025)
DyCoRM: Dynamic Criterion-Aware Reward Modeling for Text-to-Image Generation
by: Qian, Jiaying, et al.
Published: (2026)
by: Qian, Jiaying, et al.
Published: (2026)
VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
by: Xu, Haidong, et al.
Published: (2025)
by: Xu, Haidong, et al.
Published: (2025)
Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation
by: Wu, Ji-Jia, et al.
Published: (2024)
by: Wu, Ji-Jia, et al.
Published: (2024)
Hand1000: Generating Realistic Hands from Text with Only 1,000 Images
by: Zhang, Haozhuo, et al.
Published: (2024)
by: Zhang, Haozhuo, et al.
Published: (2024)
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
Text Region Multiple Information Perception Network for Scene Text Detection
by: Zheng, Jinzhi, et al.
Published: (2024)
by: Zheng, Jinzhi, et al.
Published: (2024)
Uncovering the Text Embedding in Text-to-Image Diffusion Models
by: Yu, Hu, et al.
Published: (2024)
by: Yu, Hu, et al.
Published: (2024)
MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation
by: Wei, Yuxiang, et al.
Published: (2024)
by: Wei, Yuxiang, et al.
Published: (2024)
LocInv: Localization-aware Inversion for Text-Guided Image Editing
by: Tang, Chuanming, et al.
Published: (2024)
by: Tang, Chuanming, et al.
Published: (2024)
Optimizing Prompts for Text-to-Image Generation
by: Hao, Yaru, et al.
Published: (2022)
by: Hao, Yaru, et al.
Published: (2022)
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
by: Guo, Yuxiang, et al.
Published: (2025)
by: Guo, Yuxiang, et al.
Published: (2025)
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
Text-to-Image Synthesis: A Decade Survey
by: Zhang, Nonghai, et al.
Published: (2024)
by: Zhang, Nonghai, et al.
Published: (2024)
Evaluating the Generation of Spatial Relations in Text and Image Generative Models
by: Sim, Shang Hong, et al.
Published: (2024)
by: Sim, Shang Hong, et al.
Published: (2024)
Spatially Covariant Image Registration with Text Prompts
by: Chen, Xiang, et al.
Published: (2023)
by: Chen, Xiang, et al.
Published: (2023)
Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
by: Shi, Qingyu, et al.
Published: (2025)
by: Shi, Qingyu, et al.
Published: (2025)
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
by: Zhao, Chenxi, et al.
Published: (2026)
by: Zhao, Chenxi, et al.
Published: (2026)
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
DualTSR: Unified Dual-Diffusion Transformer for Scene Text Image Super-Resolution
by: Niu, Axi, et al.
Published: (2026)
by: Niu, Axi, et al.
Published: (2026)
CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition
by: Zheng, Jinzhi, et al.
Published: (2024)
by: Zheng, Jinzhi, et al.
Published: (2024)
LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival
by: Zhao, Yuanxin, et al.
Published: (2024)
by: Zhao, Yuanxin, et al.
Published: (2024)
EPIC: Efficient Prompt Interaction for Text-Image Classification
by: Yu, Xinyao, et al.
Published: (2025)
by: Yu, Xinyao, et al.
Published: (2025)
Dynamic Prompt Optimizing for Text-to-Image Generation
by: Mo, Wenyi, et al.
Published: (2024)
by: Mo, Wenyi, et al.
Published: (2024)
Attention Calibration for Disentangled Text-to-Image Personalization
by: Zhang, Yanbing, et al.
Published: (2024)
by: Zhang, Yanbing, et al.
Published: (2024)
Similar Items
-
When Words Smile: Generating Diverse Emotional Facial Expressions from Text
by: Xu, Haidong, et al.
Published: (2024) -
Towards Text-Image Interleaved Retrieval
by: Zhang, Xin, et al.
Published: (2025) -
TIPS: Text-Image Pretraining with Spatial awareness
by: Maninis, Kevis-Kokitsi, et al.
Published: (2024) -
TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation
by: Wang, Alex Jinpeng, et al.
Published: (2025) -
Dynamic Frequency Modulation for Controllable Text-driven Image Generation
by: Shi, Tiandong, et al.
Published: (2026)