DreamText: High Fidelity Scene Text Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yibin, Zhang, Weizhong, Xu, Honghui, Jin, Cheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
High-fidelity Person-centric Subject-to-Image Synthesis
di: Wang, Yibin, et al.
Pubblicazione: (2023)
di: Wang, Yibin, et al.
Pubblicazione: (2023)
MagicFace: Training-free Universal-Style Human Image Customized Synthesis
di: Wang, Yibin, et al.
Pubblicazione: (2024)
di: Wang, Yibin, et al.
Pubblicazione: (2024)
Enhancing Object Coherence in Layout-to-Image Synthesis
di: Wang, Yibin, et al.
Pubblicazione: (2023)
di: Wang, Yibin, et al.
Pubblicazione: (2023)
PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
di: Wang, Yibin, et al.
Pubblicazione: (2024)
di: Wang, Yibin, et al.
Pubblicazione: (2024)
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
di: Zhou, Shijie, et al.
Pubblicazione: (2024)
di: Zhou, Shijie, et al.
Pubblicazione: (2024)
TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis
di: Rahman, Kazi Mahathir, et al.
Pubblicazione: (2025)
di: Rahman, Kazi Mahathir, et al.
Pubblicazione: (2025)
Plug-and-Play Multi-Concept Adaptive Blending for High-Fidelity Text-to-Image Synthesis
di: Woo, Young-Beom
Pubblicazione: (2025)
di: Woo, Young-Beom
Pubblicazione: (2025)
TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing
di: Bao, Yuchen, et al.
Pubblicazione: (2025)
di: Bao, Yuchen, et al.
Pubblicazione: (2025)
TextMamba: Scene Text Detector with Mamba
di: Zhao, Qiyan, et al.
Pubblicazione: (2025)
di: Zhao, Qiyan, et al.
Pubblicazione: (2025)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
di: Sampaio, Georgia Gabriela, et al.
Pubblicazione: (2024)
di: Sampaio, Georgia Gabriela, et al.
Pubblicazione: (2024)
Region Prompt Tuning: Fine-grained Scene Text Detection Utilizing Region Text Prompt
di: Lin, Xingtao, et al.
Pubblicazione: (2024)
di: Lin, Xingtao, et al.
Pubblicazione: (2024)
3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation
di: Zhang, Frank, et al.
Pubblicazione: (2024)
di: Zhang, Frank, et al.
Pubblicazione: (2024)
DualTSR: Unified Dual-Diffusion Transformer for Scene Text Image Super-Resolution
di: Niu, Axi, et al.
Pubblicazione: (2026)
di: Niu, Axi, et al.
Pubblicazione: (2026)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
di: Zeng, Gangyan, et al.
Pubblicazione: (2024)
di: Zeng, Gangyan, et al.
Pubblicazione: (2024)
A Generative Approach to High Fidelity 3D Reconstruction from Text Data
di: R, Venkat Kumar, et al.
Pubblicazione: (2025)
di: R, Venkat Kumar, et al.
Pubblicazione: (2025)
DreamPolisher: Towards High-Quality Text-to-3D Generation via Geometric Diffusion
di: Lin, Yuanze, et al.
Pubblicazione: (2024)
di: Lin, Yuanze, et al.
Pubblicazione: (2024)
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
di: Wang, Lizhen, et al.
Pubblicazione: (2025)
di: Wang, Lizhen, et al.
Pubblicazione: (2025)
InstructOCR: Instruction Boosting Scene Text Spotting
di: Duan, Chen, et al.
Pubblicazione: (2024)
di: Duan, Chen, et al.
Pubblicazione: (2024)
Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
di: Cao, Yushe, et al.
Pubblicazione: (2025)
di: Cao, Yushe, et al.
Pubblicazione: (2025)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
di: Wang, Xiao, et al.
Pubblicazione: (2025)
di: Wang, Xiao, et al.
Pubblicazione: (2025)
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
di: Xie, Yu, et al.
Pubblicazione: (2025)
di: Xie, Yu, et al.
Pubblicazione: (2025)
DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning
di: Ye, Fulong, et al.
Pubblicazione: (2025)
di: Ye, Fulong, et al.
Pubblicazione: (2025)
StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
di: Simonyan, Aleksandr, et al.
Pubblicazione: (2026)
di: Simonyan, Aleksandr, et al.
Pubblicazione: (2026)
HAAP: Vision-context Hierarchical Attention Autoregressive with Adaptive Permutation for Scene Text Recognition
di: Chen, Honghui, et al.
Pubblicazione: (2024)
di: Chen, Honghui, et al.
Pubblicazione: (2024)
PaintScene4D: Consistent 4D Scene Generation from Text Prompts
di: Gupta, Vinayak, et al.
Pubblicazione: (2024)
di: Gupta, Vinayak, et al.
Pubblicazione: (2024)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
di: Maeda, Koki, et al.
Pubblicazione: (2026)
di: Maeda, Koki, et al.
Pubblicazione: (2026)
QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation
di: Wasfy, Ahmed, et al.
Pubblicazione: (2025)
di: Wasfy, Ahmed, et al.
Pubblicazione: (2025)
A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2025)
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2025)
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
di: De, Anik, et al.
Pubblicazione: (2025)
di: De, Anik, et al.
Pubblicazione: (2025)
Layered Diffusion Model for One-Shot High Resolution Text-to-Image Synthesis
di: Khwaja, Emaad, et al.
Pubblicazione: (2024)
di: Khwaja, Emaad, et al.
Pubblicazione: (2024)
Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
di: Vatsa, Mayank, et al.
Pubblicazione: (2025)
di: Vatsa, Mayank, et al.
Pubblicazione: (2025)
InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
di: Wang, Cong, et al.
Pubblicazione: (2023)
di: Wang, Cong, et al.
Pubblicazione: (2023)
DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models
di: Xing, Ximing, et al.
Pubblicazione: (2023)
di: Xing, Ximing, et al.
Pubblicazione: (2023)
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
di: Qin, Can, et al.
Pubblicazione: (2024)
di: Qin, Can, et al.
Pubblicazione: (2024)
Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
di: Hu, Taihang, et al.
Pubblicazione: (2024)
di: Hu, Taihang, et al.
Pubblicazione: (2024)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
TSTMotion: Training-free Scene-aware Text-to-motion Generation
di: Guo, Ziyan, et al.
Pubblicazione: (2025)
di: Guo, Ziyan, et al.
Pubblicazione: (2025)
A Large-scale Dataset for Robust Complex Anime Scene Text Detection
di: Dong, Ziyi, et al.
Pubblicazione: (2025)
di: Dong, Ziyi, et al.
Pubblicazione: (2025)
BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving
di: Tang, Tao, et al.
Pubblicazione: (2024)
di: Tang, Tao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
High-fidelity Person-centric Subject-to-Image Synthesis
di: Wang, Yibin, et al.
Pubblicazione: (2023) -
MagicFace: Training-free Universal-Style Human Image Customized Synthesis
di: Wang, Yibin, et al.
Pubblicazione: (2024) -
Enhancing Object Coherence in Layout-to-Image Synthesis
di: Wang, Yibin, et al.
Pubblicazione: (2023) -
PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
di: Wang, Yibin, et al.
Pubblicazione: (2024) -
DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
di: Zhou, Shijie, et al.
Pubblicazione: (2024)