Beyond Flat Text: Dual Self-inherited Guidance for Visual Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Minxing, Xia, Zixun, Chen, Liaojun, Li, Zhenhang, Zeng, Weichao, Wang, Jianye, Cheng, Wentao, Wang, Yaxing, Zhou, Yu, Yang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
by: Zeng, Weichao, et al.
Published: (2024)
by: Zeng, Weichao, et al.
Published: (2024)
Visual Text Meets Low-level Vision: A Comprehensive Survey on Visual Text Processing
by: Shu, Yan, et al.
Published: (2024)
by: Shu, Yan, et al.
Published: (2024)
First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
by: Li, Zhenhang, et al.
Published: (2024)
by: Li, Zhenhang, et al.
Published: (2024)
Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance
by: Luo, Minxing, et al.
Published: (2025)
by: Luo, Minxing, et al.
Published: (2025)
Visual Text Processing: A Comprehensive Review and Unified Evaluation
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
RotationDrag: Point-based Image Editing with Rotated Diffusion Features
by: Luo, Minxing, et al.
Published: (2024)
by: Luo, Minxing, et al.
Published: (2024)
TADoc: Robust Time-Aware Document Image Dewarping
by: Zhao, Fangmin, et al.
Published: (2025)
by: Zhao, Fangmin, et al.
Published: (2025)
Encoding Urban Ecologies: Automated Building Archetype Generation through Self-Supervised Learning for Energy Modeling
by: Zhuang, Xinwei, et al.
Published: (2024)
by: Zhuang, Xinwei, et al.
Published: (2024)
Dual-Stream Diffusion Net for Text-to-Video Generation
by: Liu, Binhui, et al.
Published: (2023)
by: Liu, Binhui, et al.
Published: (2023)
Conditional Text-to-Image Generation with Reference Guidance
by: Kim, Taewook, et al.
Published: (2024)
by: Kim, Taewook, et al.
Published: (2024)
The Devil is in Fine-tuning and Long-tailed Problems:A New Benchmark for Scene Text Detection
by: Cao, Tianjiao, et al.
Published: (2025)
by: Cao, Tianjiao, et al.
Published: (2025)
Uni-DocDiff: A Unified Document Restoration Model Based on Diffusion
by: Zhao, Fangmin, et al.
Published: (2025)
by: Zhao, Fangmin, et al.
Published: (2025)
Text‐to‐3D City: Plan‐then‐Execute Urban Generation With LLM Planners and Procedural Synthesis
by: Xiaohang Dong, et al.
Published: (2026)
by: Xiaohang Dong, et al.
Published: (2026)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
DTVI: Dual-Stage Textual and Visual Intervention for Safe Text-to-Image Generation
by: Tan, Binhong, et al.
Published: (2026)
by: Tan, Binhong, et al.
Published: (2026)
Dual Orthogonal Guidance for Robust Diffusion-based Handwritten Text Generation
by: Nikolaidou, Konstantina, et al.
Published: (2025)
by: Nikolaidou, Konstantina, et al.
Published: (2025)
IMTBench: A Multi-Scenario Cross-Modal Collaborative Evaluation Benchmark for In-Image Machine Translation
by: Lyu, Jiahao, et al.
Published: (2026)
by: Lyu, Jiahao, et al.
Published: (2026)
Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation
by: Xie, Dian, et al.
Published: (2026)
by: Xie, Dian, et al.
Published: (2026)
Investigating Text Insulation and Attention Mechanisms for Complex Visual Text Generation
by: Tai, Ying, et al.
Published: (2025)
by: Tai, Ying, et al.
Published: (2025)
Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation Ability
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
Mustango: Toward Controllable Text-to-Music Generation
by: Melechovsky, Jan, et al.
Published: (2023)
by: Melechovsky, Jan, et al.
Published: (2023)
Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects
by: Qiu, Weimin, et al.
Published: (2024)
by: Qiu, Weimin, et al.
Published: (2024)
CO Adsorption Energy Match at Dual Sites Drives C–C Coupling Activity in Rare Earth–Cu CO 2 Reduction Catalysts
by: Minxing Shu, et al.
Published: (2026)
by: Minxing Shu, et al.
Published: (2026)
Text-Animator: Controllable Visual Text Video Generation
by: Liu, Lin, et al.
Published: (2024)
by: Liu, Lin, et al.
Published: (2024)
Self-Explaining Hypergraph Neural Networks for Diagnosis Prediction
by: Yu, Leisheng, et al.
Published: (2025)
by: Yu, Leisheng, et al.
Published: (2025)
Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
The Role of Video Generation in Enhancing Data-Limited Action Understanding
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Visual Text Generation in the Wild
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
DICE: Distilling Classifier-Free Guidance into Text Embeddings
by: Zhou, Zhenyu, et al.
Published: (2025)
by: Zhou, Zhenyu, et al.
Published: (2025)
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
StyleDiffusion: Prompt-Embedding Inversion for Text-Based Editing
by: Li, Senmao, et al.
Published: (2023)
by: Li, Senmao, et al.
Published: (2023)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
by: Hu, Taihang, et al.
Published: (2024)
by: Hu, Taihang, et al.
Published: (2024)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
by: Tong, Chengzhuo, et al.
Published: (2026)
by: Tong, Chengzhuo, et al.
Published: (2026)
GAP: Gaussianize Any Point Clouds with Text Guidance
by: Zhang, Weiqi, et al.
Published: (2025)
by: Zhang, Weiqi, et al.
Published: (2025)
How is Visual Attention Influenced by Text Guidance? Database and Model
by: Sun, Yinan, et al.
Published: (2024)
by: Sun, Yinan, et al.
Published: (2024)
Zero-Shot Visual Concept Blending Without Text Guidance
by: Makino, Hiroya, et al.
Published: (2025)
by: Makino, Hiroya, et al.
Published: (2025)
TextMaster: A Unified Framework for Realistic Text Editing via Glyph-Style Dual-Control
by: Yan, Zhenyu, et al.
Published: (2024)
by: Yan, Zhenyu, et al.
Published: (2024)
SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance
by: Shen, Guibao, et al.
Published: (2024)
by: Shen, Guibao, et al.
Published: (2024)
Dual Refinement Cycle Learning: Unsupervised Text Classification of Mamba and Community Detection on Text Attributed Graph
by: Wang, Hong, et al.
Published: (2025)
by: Wang, Hong, et al.
Published: (2025)
Similar Items
-
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
by: Zeng, Weichao, et al.
Published: (2024) -
Visual Text Meets Low-level Vision: A Comprehensive Survey on Visual Text Processing
by: Shu, Yan, et al.
Published: (2024) -
First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
by: Li, Zhenhang, et al.
Published: (2024) -
Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance
by: Luo, Minxing, et al.
Published: (2025) -
Visual Text Processing: A Comprehensive Review and Unified Evaluation
by: Shu, Yan, et al.
Published: (2025)