DreamText: High Fidelity Scene Text Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yibin, Zhang, Weizhong, Xu, Honghui, Jin, Cheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
High-fidelity Person-centric Subject-to-Image Synthesis
by: Wang, Yibin, et al.
Published: (2023)
by: Wang, Yibin, et al.
Published: (2023)
MagicFace: Training-free Universal-Style Human Image Customized Synthesis
by: Wang, Yibin, et al.
Published: (2024)
by: Wang, Yibin, et al.
Published: (2024)
Enhancing Object Coherence in Layout-to-Image Synthesis
by: Wang, Yibin, et al.
Published: (2023)
by: Wang, Yibin, et al.
Published: (2023)
PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
by: Wang, Yibin, et al.
Published: (2024)
by: Wang, Yibin, et al.
Published: (2024)
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis
by: Rahman, Kazi Mahathir, et al.
Published: (2025)
by: Rahman, Kazi Mahathir, et al.
Published: (2025)
Plug-and-Play Multi-Concept Adaptive Blending for High-Fidelity Text-to-Image Synthesis
by: Woo, Young-Beom
Published: (2025)
by: Woo, Young-Beom
Published: (2025)
TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing
by: Bao, Yuchen, et al.
Published: (2025)
by: Bao, Yuchen, et al.
Published: (2025)
TextMamba: Scene Text Detector with Mamba
by: Zhao, Qiyan, et al.
Published: (2025)
by: Zhao, Qiyan, et al.
Published: (2025)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
Region Prompt Tuning: Fine-grained Scene Text Detection Utilizing Region Text Prompt
by: Lin, Xingtao, et al.
Published: (2024)
by: Lin, Xingtao, et al.
Published: (2024)
3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation
by: Zhang, Frank, et al.
Published: (2024)
by: Zhang, Frank, et al.
Published: (2024)
DualTSR: Unified Dual-Diffusion Transformer for Scene Text Image Super-Resolution
by: Niu, Axi, et al.
Published: (2026)
by: Niu, Axi, et al.
Published: (2026)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
by: Zeng, Gangyan, et al.
Published: (2024)
by: Zeng, Gangyan, et al.
Published: (2024)
A Generative Approach to High Fidelity 3D Reconstruction from Text Data
by: R, Venkat Kumar, et al.
Published: (2025)
by: R, Venkat Kumar, et al.
Published: (2025)
DreamPolisher: Towards High-Quality Text-to-3D Generation via Geometric Diffusion
by: Lin, Yuanze, et al.
Published: (2024)
by: Lin, Yuanze, et al.
Published: (2024)
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
by: Wang, Lizhen, et al.
Published: (2025)
by: Wang, Lizhen, et al.
Published: (2025)
InstructOCR: Instruction Boosting Scene Text Spotting
by: Duan, Chen, et al.
Published: (2024)
by: Duan, Chen, et al.
Published: (2024)
Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
by: Cao, Yushe, et al.
Published: (2025)
by: Cao, Yushe, et al.
Published: (2025)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
by: Xie, Yu, et al.
Published: (2025)
by: Xie, Yu, et al.
Published: (2025)
DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning
by: Ye, Fulong, et al.
Published: (2025)
by: Ye, Fulong, et al.
Published: (2025)
StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
by: Simonyan, Aleksandr, et al.
Published: (2026)
by: Simonyan, Aleksandr, et al.
Published: (2026)
HAAP: Vision-context Hierarchical Attention Autoregressive with Adaptive Permutation for Scene Text Recognition
by: Chen, Honghui, et al.
Published: (2024)
by: Chen, Honghui, et al.
Published: (2024)
PaintScene4D: Consistent 4D Scene Generation from Text Prompts
by: Gupta, Vinayak, et al.
Published: (2024)
by: Gupta, Vinayak, et al.
Published: (2024)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
by: Maeda, Koki, et al.
Published: (2026)
by: Maeda, Koki, et al.
Published: (2026)
QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation
by: Wasfy, Ahmed, et al.
Published: (2025)
by: Wasfy, Ahmed, et al.
Published: (2025)
A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
by: De, Anik, et al.
Published: (2025)
by: De, Anik, et al.
Published: (2025)
Layered Diffusion Model for One-Shot High Resolution Text-to-Image Synthesis
by: Khwaja, Emaad, et al.
Published: (2024)
by: Khwaja, Emaad, et al.
Published: (2024)
Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
by: Vatsa, Mayank, et al.
Published: (2025)
by: Vatsa, Mayank, et al.
Published: (2025)
InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
by: Cai, Xinhao, et al.
Published: (2025)
by: Cai, Xinhao, et al.
Published: (2025)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
by: Wang, Cong, et al.
Published: (2023)
by: Wang, Cong, et al.
Published: (2023)
DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models
by: Xing, Ximing, et al.
Published: (2023)
by: Xing, Ximing, et al.
Published: (2023)
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
by: Qin, Can, et al.
Published: (2024)
by: Qin, Can, et al.
Published: (2024)
Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
by: Hu, Taihang, et al.
Published: (2024)
by: Hu, Taihang, et al.
Published: (2024)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
by: Lee, Wonjun, et al.
Published: (2025)
by: Lee, Wonjun, et al.
Published: (2025)
TSTMotion: Training-free Scene-aware Text-to-motion Generation
by: Guo, Ziyan, et al.
Published: (2025)
by: Guo, Ziyan, et al.
Published: (2025)
A Large-scale Dataset for Robust Complex Anime Scene Text Detection
by: Dong, Ziyi, et al.
Published: (2025)
by: Dong, Ziyi, et al.
Published: (2025)
BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving
by: Tang, Tao, et al.
Published: (2024)
by: Tang, Tao, et al.
Published: (2024)
Similar Items
-
High-fidelity Person-centric Subject-to-Image Synthesis
by: Wang, Yibin, et al.
Published: (2023) -
MagicFace: Training-free Universal-Style Human Image Customized Synthesis
by: Wang, Yibin, et al.
Published: (2024) -
Enhancing Object Coherence in Layout-to-Image Synthesis
by: Wang, Yibin, et al.
Published: (2023) -
PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
by: Wang, Yibin, et al.
Published: (2024) -
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
by: Zhou, Shijie, et al.
Published: (2024)