WordCon: Word-level Typography Control in Scene Text Rendering
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Wenda, Song, Yiren, Rao, Zihan, Zhang, Dengming, Liu, Jiaming, Zou, Xingxing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FonTS: Text Rendering with Typography and Style Controls
by: Shi, Wenda, et al.
Published: (2024)
by: Shi, Wenda, et al.
Published: (2024)
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
by: Lu, Runnan, et al.
Published: (2025)
by: Lu, Runnan, et al.
Published: (2025)
AnySurf: Any Surface Generation with Directed Edge
by: Shi, Wenda, et al.
Published: (2026)
by: Shi, Wenda, et al.
Published: (2026)
WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending
by: Wang, Zhe, et al.
Published: (2025)
by: Wang, Zhe, et al.
Published: (2025)
EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer
by: Zhang, Yuxuan, et al.
Published: (2025)
by: Zhang, Yuxuan, et al.
Published: (2025)
WordArt Designer API: User-Driven Artistic Typography Synthesis with Large Language Models on ModelScope
by: He, Jun-Yan, et al.
Published: (2024)
by: He, Jun-Yan, et al.
Published: (2024)
Loom: Diffusion-Transformer for Interleaved Generation
by: Ye, Mingcheng, et al.
Published: (2025)
by: Ye, Mingcheng, et al.
Published: (2025)
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
by: Zhang, Yunzhi, et al.
Published: (2024)
by: Zhang, Yunzhi, et al.
Published: (2024)
Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs
by: Zheng, Yiren, et al.
Published: (2026)
by: Zheng, Yiren, et al.
Published: (2026)
Dynamic Typography: Bringing Text to Life via Video Diffusion Prior
by: Liu, Zichen, et al.
Published: (2024)
by: Liu, Zichen, et al.
Published: (2024)
Stable-Makeup: When Real-World Makeup Transfer Meets Diffusion Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
by: Cui, Mingxuan, et al.
Published: (2026)
by: Cui, Mingxuan, et al.
Published: (2026)
Multiplane Prior Guided Few-Shot Aerial Scene Rendering
by: Gao, Zihan, et al.
Published: (2024)
by: Gao, Zihan, et al.
Published: (2024)
Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography
by: Chung, Nhat, et al.
Published: (2024)
by: Chung, Nhat, et al.
Published: (2024)
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)
by: D'Incà, Moreno, et al.
Published: (2024)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
by: Huang, Mengqi, et al.
Published: (2024)
by: Huang, Mengqi, et al.
Published: (2024)
Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound
by: Zhang, Dengming, et al.
Published: (2025)
by: Zhang, Dengming, et al.
Published: (2025)
Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line Generation
by: Dai, Gang, et al.
Published: (2025)
by: Dai, Gang, et al.
Published: (2025)
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
by: Lu, Jianglin, et al.
Published: (2026)
by: Lu, Jianglin, et al.
Published: (2026)
Learning Continuous 3D Words for Text-to-Image Generation
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
by: Luo, Zhi, et al.
Published: (2025)
by: Luo, Zhi, et al.
Published: (2025)
High-quality Pseudo-labeling for Point Cloud Segmentation with Scene-level Annotation
by: Duan, Lunhao, et al.
Published: (2025)
by: Duan, Lunhao, et al.
Published: (2025)
LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Rendering and Control
by: Qu, Delin, et al.
Published: (2024)
by: Qu, Delin, et al.
Published: (2024)
Stable-Hair: Real-World Hair Transfer via Diffusion Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
by: Guo, Hailong, et al.
Published: (2025)
by: Guo, Hailong, et al.
Published: (2025)
FlexDrive: Toward Trajectory Flexibility in Driving Scene Reconstruction and Rendering
by: Zhou, Jingqiu, et al.
Published: (2025)
by: Zhou, Jingqiu, et al.
Published: (2025)
Reflections Unlock: Geometry-Aware Reflection Disentanglement in 3D Gaussian Splatting for Photorealistic Scenes Rendering
by: Song, Jiayi, et al.
Published: (2025)
by: Song, Jiayi, et al.
Published: (2025)
Typography-Based Monocular Distance Estimation Framework for Vehicle Safety Systems
by: Reddy, Manognya Lokesh, et al.
Published: (2026)
by: Reddy, Manognya Lokesh, et al.
Published: (2026)
Kinetic Typography Diffusion Model
by: Park, Seonmi, et al.
Published: (2024)
by: Park, Seonmi, et al.
Published: (2024)
Intelligent Artistic Typography: A Comprehensive Review of Artistic Text Design and Generation
by: Bai, Yuhang, et al.
Published: (2024)
by: Bai, Yuhang, et al.
Published: (2024)
NCRF: Neural Contact Radiance Fields for Free-Viewpoint Rendering of Hand-Object Interaction
by: Zhang, Zhongqun, et al.
Published: (2024)
by: Zhang, Zhongqun, et al.
Published: (2024)
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
Beautiful Images, Toxic Words: Understanding and Addressing Offensive Text in Generated Images
by: Kumar, Aditya, et al.
Published: (2025)
by: Kumar, Aditya, et al.
Published: (2025)
Attention vs LSTM: Improving Word-level BISINDO Recognition
by: Kautsar, Muchammad Daniyal, et al.
Published: (2024)
by: Kautsar, Muchammad Daniyal, et al.
Published: (2024)
GIR: 3D Gaussian Inverse Rendering for Relightable Scene Factorization
by: Shi, Yahao, et al.
Published: (2023)
by: Shi, Yahao, et al.
Published: (2023)
NieR: Normal-Based Lighting Scene Rendering
by: Wang, Hongsheng, et al.
Published: (2024)
by: Wang, Hongsheng, et al.
Published: (2024)
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
WordRobe: Text-Guided Generation of Textured 3D Garments
by: Srivastava, Astitva, et al.
Published: (2024)
by: Srivastava, Astitva, et al.
Published: (2024)
RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
by: Pan, Mingjie, et al.
Published: (2023)
by: Pan, Mingjie, et al.
Published: (2023)
Fast Personalized Text-to-Image Syntheses With Attention Injection
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
Similar Items
-
FonTS: Text Rendering with Typography and Style Controls
by: Shi, Wenda, et al.
Published: (2024) -
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
by: Lu, Runnan, et al.
Published: (2025) -
AnySurf: Any Surface Generation with Directed Edge
by: Shi, Wenda, et al.
Published: (2026) -
WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending
by: Wang, Zhe, et al.
Published: (2025) -
EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer
by: Zhang, Yuxuan, et al.
Published: (2025)