HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhuang, Shuhan, Huang, Mengqi, Fu, Fengyi, Chen, Nan, Lei, Bohan, Mao, Zhendong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer Learning
por: Fu, Fengyi, et al.
Publicado: (2025)
por: Fu, Fengyi, et al.
Publicado: (2025)
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
por: Lin, Yijing, et al.
Publicado: (2025)
por: Lin, Yijing, et al.
Publicado: (2025)
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
por: Chen, Nan, et al.
Publicado: (2025)
por: Chen, Nan, et al.
Publicado: (2025)
D$^2$iT: Dynamic Diffusion Transformer for Accurate Image Generation
por: Jia, Weinan, et al.
Publicado: (2025)
por: Jia, Weinan, et al.
Publicado: (2025)
GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering
por: Shuai, Xincheng, et al.
Publicado: (2026)
por: Shuai, Xincheng, et al.
Publicado: (2026)
Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering
por: Liu, Zeyu, et al.
Publicado: (2024)
por: Liu, Zeyu, et al.
Publicado: (2024)
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
por: Chen, Nan, et al.
Publicado: (2024)
por: Chen, Nan, et al.
Publicado: (2024)
Glyph-ByT5-v2: A Strong Aesthetic Baseline for Accurate Multilingual Visual Text Rendering
por: Liu, Zeyu, et al.
Publicado: (2024)
por: Liu, Zeyu, et al.
Publicado: (2024)
FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection
por: Zhang, Ruiqiang, et al.
Publicado: (2026)
por: Zhang, Ruiqiang, et al.
Publicado: (2026)
GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows
por: Yan, Zexuan, et al.
Publicado: (2026)
por: Yan, Zexuan, et al.
Publicado: (2026)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
por: Huang, Mengqi, et al.
Publicado: (2024)
por: Huang, Mengqi, et al.
Publicado: (2024)
DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image Generation
por: Huang, Mengqi, et al.
Publicado: (2022)
por: Huang, Mengqi, et al.
Publicado: (2022)
UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis
por: Wang, Yuanrui, et al.
Publicado: (2025)
por: Wang, Yuanrui, et al.
Publicado: (2025)
Training-Free Occluded Text Rendering via Glyph Priors and Attention-Guided Semantic Blending
por: Hou, Jingqi, et al.
Publicado: (2026)
por: Hou, Jingqi, et al.
Publicado: (2026)
Sentiment-oriented Transformer-based Variational Autoencoder Network for Live Video Commenting
por: Fu, Fengyi, et al.
Publicado: (2024)
por: Fu, Fengyi, et al.
Publicado: (2024)
GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing
por: Wang, Tong, et al.
Publicado: (2025)
por: Wang, Tong, et al.
Publicado: (2025)
GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language Models
por: Ma, Jian, et al.
Publicado: (2024)
por: Ma, Jian, et al.
Publicado: (2024)
Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute Editing
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
NativeTok: Native Visual Tokenization for Improved Image Generation
por: Wu, Bin, et al.
Publicado: (2026)
por: Wu, Bin, et al.
Publicado: (2026)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
por: Jia, Weinan, et al.
Publicado: (2025)
por: Jia, Weinan, et al.
Publicado: (2025)
DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization
por: Wang, Wenchuan, et al.
Publicado: (2025)
por: Wang, Wenchuan, et al.
Publicado: (2025)
LOGO: Video Text Spotting with Language Collaboration and Glyph Perception Model
por: Liu, Hongen, et al.
Publicado: (2024)
por: Liu, Hongen, et al.
Publicado: (2024)
Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware Training
por: Li, Wenbo, et al.
Publicado: (2024)
por: Li, Wenbo, et al.
Publicado: (2024)
TextMaster: A Unified Framework for Realistic Text Editing via Glyph-Style Dual-Control
por: Yan, Zhenyu, et al.
Publicado: (2024)
por: Yan, Zhenyu, et al.
Publicado: (2024)
DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
por: Pan, Wei, et al.
Publicado: (2025)
por: Pan, Wei, et al.
Publicado: (2025)
Disentangling Hardness from Noise: An Uncertainty-Driven Model-Agnostic Framework for Long-Tailed Remote Sensing Classification
por: Ding, Chi, et al.
Publicado: (2026)
por: Ding, Chi, et al.
Publicado: (2026)
Lance: Unified Multimodal Modeling by Multi-Task Synergy
por: Fu, Fengyi, et al.
Publicado: (2026)
por: Fu, Fengyi, et al.
Publicado: (2026)
Hierarchical Context Alignment with Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction
por: Li, Bohan, et al.
Publicado: (2024)
por: Li, Bohan, et al.
Publicado: (2024)
AnyArtisticGlyph: Multilingual Controllable Artistic Glyph Generation
por: Lu, Xiongbo, et al.
Publicado: (2025)
por: Lu, Xiongbo, et al.
Publicado: (2025)
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
por: Zhang, Hui, et al.
Publicado: (2026)
por: Zhang, Hui, et al.
Publicado: (2026)
Streaming Video Diffusion: Online Video Editing with Diffusion Models
por: Chen, Feng, et al.
Publicado: (2024)
por: Chen, Feng, et al.
Publicado: (2024)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
por: Chen, Zhuokun, et al.
Publicado: (2026)
por: Chen, Zhuokun, et al.
Publicado: (2026)
Text-guided Foundation Model Adaptation for Long-Tailed Medical Image Classification
por: Li, Sirui, et al.
Publicado: (2024)
por: Li, Sirui, et al.
Publicado: (2024)
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
por: Cui, Mingxuan, et al.
Publicado: (2026)
por: Cui, Mingxuan, et al.
Publicado: (2026)
VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
por: Wu, Shaojin, et al.
Publicado: (2024)
por: Wu, Shaojin, et al.
Publicado: (2024)
TextPixs: Glyph-Conditioned Diffusion with Character-Aware Attention and OCR-Guided Supervision
por: Gillani, Syeda Anshrah, et al.
Publicado: (2025)
por: Gillani, Syeda Anshrah, et al.
Publicado: (2025)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
por: Mao, Weian, et al.
Publicado: (2026)
por: Mao, Weian, et al.
Publicado: (2026)
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
por: Lu, Runnan, et al.
Publicado: (2025)
por: Lu, Runnan, et al.
Publicado: (2025)
Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal Denoiser
por: Cai, Qingyuan, et al.
Publicado: (2024)
por: Cai, Qingyuan, et al.
Publicado: (2024)
LongVLM: Efficient Long Video Understanding via Large Language Models
por: Weng, Yuetian, et al.
Publicado: (2024)
por: Weng, Yuetian, et al.
Publicado: (2024)
Ejemplares similares
-
LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer Learning
por: Fu, Fengyi, et al.
Publicado: (2025) -
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
por: Lin, Yijing, et al.
Publicado: (2025) -
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
por: Chen, Nan, et al.
Publicado: (2025) -
D$^2$iT: Dynamic Diffusion Transformer for Accurate Image Generation
por: Jia, Weinan, et al.
Publicado: (2025) -
GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering
por: Shuai, Xincheng, et al.
Publicado: (2026)