Visual Text Meets Low-level Vision: A Comprehensive Survey on Visual Text Processing
Fuente:
arXiv
Saved in:
| Main Authors: | Shu, Yan, Zeng, Weichao, Li, Zhenhang, Zhao, Fangmin, Zhou, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Text Processing: A Comprehensive Review and Unified Evaluation
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
by: Li, Zhenhang, et al.
Published: (2024)
by: Li, Zhenhang, et al.
Published: (2024)
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
by: Zeng, Weichao, et al.
Published: (2024)
by: Zeng, Weichao, et al.
Published: (2024)
TADoc: Robust Time-Aware Document Image Dewarping
by: Zhao, Fangmin, et al.
Published: (2025)
by: Zhao, Fangmin, et al.
Published: (2025)
Beyond Flat Text: Dual Self-inherited Guidance for Visual Text Generation
by: Luo, Minxing, et al.
Published: (2025)
by: Luo, Minxing, et al.
Published: (2025)
Mostly Text, Smart Visuals: Asymmetric Text-Visual Pruning for Large Vision-Language Models
by: Li, Sijie, et al.
Published: (2026)
by: Li, Sijie, et al.
Published: (2026)
Uni-DocDiff: A Unified Document Restoration Model Based on Diffusion
by: Zhao, Fangmin, et al.
Published: (2025)
by: Zhao, Fangmin, et al.
Published: (2025)
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
by: Tang, Jingqun, et al.
Published: (2024)
by: Tang, Jingqun, et al.
Published: (2024)
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
by: Berasi, Davide, et al.
Published: (2025)
by: Berasi, Davide, et al.
Published: (2025)
FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
by: Bo, Weihao, et al.
Published: (2025)
by: Bo, Weihao, et al.
Published: (2025)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
by: Huang, Chen, et al.
Published: (2026)
by: Huang, Chen, et al.
Published: (2026)
Vision Learners Meet Web Image-Text Pairs
by: Zhao, Bingchen, et al.
Published: (2023)
by: Zhao, Bingchen, et al.
Published: (2023)
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
by: Lan, Tian, et al.
Published: (2025)
by: Lan, Tian, et al.
Published: (2025)
Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
by: Wang, Ruiyu, et al.
Published: (2025)
by: Wang, Ruiyu, et al.
Published: (2025)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
Describe Anything Model for Visual Question Answering on Text-rich Images
by: Vu, Yen-Linh, et al.
Published: (2025)
by: Vu, Yen-Linh, et al.
Published: (2025)
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2024)
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2024)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
Harmonizing Visual Text Comprehension and Generation
by: Zhao, Zhen, et al.
Published: (2024)
by: Zhao, Zhen, et al.
Published: (2024)
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
TSP3D: Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding
by: Guo, Wenxuan, et al.
Published: (2025)
by: Guo, Wenxuan, et al.
Published: (2025)
Towards Visual Text Grounding of Multimodal Large Language Model
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
by: Maini, Pratyush, et al.
Published: (2023)
by: Maini, Pratyush, et al.
Published: (2023)
Glyph: Scaling Context Windows via Visual-Text Compression
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
Learning from Imperfect Text Guidance: Robust Long-Tail Visual Recognition with High-Noise Label
by: Li, Mengke, et al.
Published: (2026)
by: Li, Mengke, et al.
Published: (2026)
V-FAT: Benchmarking Visual Fidelity Against Text-bias
by: Wang, Ziteng, et al.
Published: (2026)
by: Wang, Ziteng, et al.
Published: (2026)
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models
by: Clavié, Benjamin, et al.
Published: (2025)
by: Clavié, Benjamin, et al.
Published: (2025)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
by: Yang, Zhoufaran, et al.
Published: (2025)
by: Yang, Zhoufaran, et al.
Published: (2025)
CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data
by: Liu, Disheng, et al.
Published: (2025)
by: Liu, Disheng, et al.
Published: (2025)
Visual Prompting in Multimodal Large Language Models: A Survey
by: Wu, Junda, et al.
Published: (2024)
by: Wu, Junda, et al.
Published: (2024)
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
by: Liu, Zhuchenyang, et al.
Published: (2026)
by: Liu, Zhuchenyang, et al.
Published: (2026)
A Comprehensive Survey on Visual Concept Mining in Text-to-image Diffusion Models
by: Li, Ziqiang, et al.
Published: (2025)
by: Li, Ziqiang, et al.
Published: (2025)
VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
by: Hou, Haowen, et al.
Published: (2024)
by: Hou, Haowen, et al.
Published: (2024)
WriteViT: Handwritten Text Generation with Vision Transformer
by: Nam, Dang Hoai, et al.
Published: (2025)
by: Nam, Dang Hoai, et al.
Published: (2025)
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
by: Xiao, Lei, et al.
Published: (2025)
by: Xiao, Lei, et al.
Published: (2025)
Personalized Vision via Visual In-Context Learning
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
Expressive Text-to-Image Generation with Rich Text
by: Ge, Songwei, et al.
Published: (2023)
by: Ge, Songwei, et al.
Published: (2023)
Similar Items
-
Visual Text Processing: A Comprehensive Review and Unified Evaluation
by: Shu, Yan, et al.
Published: (2025) -
First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
by: Li, Zhenhang, et al.
Published: (2024) -
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
by: Zeng, Weichao, et al.
Published: (2024) -
TADoc: Robust Time-Aware Document Image Dewarping
by: Zhao, Fangmin, et al.
Published: (2025) -
Beyond Flat Text: Dual Self-inherited Guidance for Visual Text Generation
by: Luo, Minxing, et al.
Published: (2025)