Visual Text Processing: A Comprehensive Review and Unified Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shu, Yan, Zeng, Weichao, Zhao, Fangmin, Chen, Zeyu, Li, Zhenhang, Yang, Xiaomeng, Zhou, Yu, Rota, Paolo, Bai, Xiang, Jin, Lianwen, Yin, Xu-Cheng, Sebe, Nicu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual Text Meets Low-level Vision: A Comprehensive Survey on Visual Text Processing
von: Shu, Yan, et al.
Veröffentlicht: (2024)
von: Shu, Yan, et al.
Veröffentlicht: (2024)
First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
von: Li, Zhenhang, et al.
Veröffentlicht: (2024)
von: Li, Zhenhang, et al.
Veröffentlicht: (2024)
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
von: Shu, Yan, et al.
Veröffentlicht: (2026)
von: Shu, Yan, et al.
Veröffentlicht: (2026)
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
von: Zeng, Weichao, et al.
Veröffentlicht: (2024)
von: Zeng, Weichao, et al.
Veröffentlicht: (2024)
TADoc: Robust Time-Aware Document Image Dewarping
von: Zhao, Fangmin, et al.
Veröffentlicht: (2025)
von: Zhao, Fangmin, et al.
Veröffentlicht: (2025)
Uni-DocDiff: A Unified Document Restoration Model Based on Diffusion
von: Zhao, Fangmin, et al.
Veröffentlicht: (2025)
von: Zhao, Fangmin, et al.
Veröffentlicht: (2025)
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
von: Shu, Yan, et al.
Veröffentlicht: (2025)
von: Shu, Yan, et al.
Veröffentlicht: (2025)
Beyond Flat Text: Dual Self-inherited Guidance for Visual Text Generation
von: Luo, Minxing, et al.
Veröffentlicht: (2025)
von: Luo, Minxing, et al.
Veröffentlicht: (2025)
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
Asymmetric GANs for Image-to-Image Translation
von: Tang, Hao, et al.
Veröffentlicht: (2019)
von: Tang, Hao, et al.
Veröffentlicht: (2019)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
Vision+X: A Survey on Multimodal Learning in the Light of Data
von: Zhu, Ye, et al.
Veröffentlicht: (2022)
von: Zhu, Ye, et al.
Veröffentlicht: (2022)
Progressive Evolution from Single-Point to Polygon for Scene Text
von: Deng, Linger, et al.
Veröffentlicht: (2023)
von: Deng, Linger, et al.
Veröffentlicht: (2023)
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
von: Shu, Yan, et al.
Veröffentlicht: (2025)
von: Shu, Yan, et al.
Veröffentlicht: (2025)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
von: Xing, Songlong, et al.
Veröffentlicht: (2025)
von: Xing, Songlong, et al.
Veröffentlicht: (2025)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
von: Song, Yue, et al.
Veröffentlicht: (2023)
von: Song, Yue, et al.
Veröffentlicht: (2023)
Hyperbolic Busemann Neural Networks
von: Chen, Ziheng, et al.
Veröffentlicht: (2026)
von: Chen, Ziheng, et al.
Veröffentlicht: (2026)
URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding
von: Shi, Yongxin, et al.
Veröffentlicht: (2025)
von: Shi, Yongxin, et al.
Veröffentlicht: (2025)
Textual Knowledge Matters: Cross-Modality Co-Teaching for Generalized Visual Class Discovery
von: Zheng, Haiyang, et al.
Veröffentlicht: (2024)
von: Zheng, Haiyang, et al.
Veröffentlicht: (2024)
ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
von: Rigo, Andrea, et al.
Veröffentlicht: (2025)
von: Rigo, Andrea, et al.
Veröffentlicht: (2025)
LEGO: Self-Supervised Representation Learning for Scene Text Images
von: Ren, Yujin, et al.
Veröffentlicht: (2024)
von: Ren, Yujin, et al.
Veröffentlicht: (2024)
Reverse Personalization
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)
von: Kung, Han-Wei, et al.
Veröffentlicht: (2025)
Loomis Painter: Reconstructing the Painting Process
von: Pobitzer, Markus, et al.
Veröffentlicht: (2025)
von: Pobitzer, Markus, et al.
Veröffentlicht: (2025)
Deep Learning-Based Object Pose Estimation: A Comprehensive Survey
von: Liu, Jian, et al.
Veröffentlicht: (2024)
von: Liu, Jian, et al.
Veröffentlicht: (2024)
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
von: Zuo, Zhi, et al.
Veröffentlicht: (2025)
von: Zuo, Zhi, et al.
Veröffentlicht: (2025)
Democratizing Fine-grained Visual Recognition with Large Language Models
von: Liu, Mingxuan, et al.
Veröffentlicht: (2024)
von: Liu, Mingxuan, et al.
Veröffentlicht: (2024)
Omni-IML: Towards Unified Image Manipulation Localization
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
Curriculum-DPO++: Direct Preference Optimization via Data and Model Curricula for Text-to-Image Generation
von: Croitoru, Florinel-Alin, et al.
Veröffentlicht: (2026)
von: Croitoru, Florinel-Alin, et al.
Veröffentlicht: (2026)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
von: Li, Chenxi, et al.
Veröffentlicht: (2025)
von: Li, Chenxi, et al.
Veröffentlicht: (2025)
3D Part Segmentation via Geometric Aggregation of 2D Visual Features
von: Garosi, Marco, et al.
Veröffentlicht: (2024)
von: Garosi, Marco, et al.
Veröffentlicht: (2024)
DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
von: Pan, Wei, et al.
Veröffentlicht: (2025)
von: Pan, Wei, et al.
Veröffentlicht: (2025)
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
von: D'Incà, Moreno, et al.
Veröffentlicht: (2024)
von: D'Incà, Moreno, et al.
Veröffentlicht: (2024)
Anti-Forgetting Adaptation for Unsupervised Person Re-identification
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Riemannian Networks over Full-Rank Correlation Matrices
von: Chen, Ziheng, et al.
Veröffentlicht: (2026)
von: Chen, Ziheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Visual Text Meets Low-level Vision: A Comprehensive Survey on Visual Text Processing
von: Shu, Yan, et al.
Veröffentlicht: (2024) -
First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
von: Li, Zhenhang, et al.
Veröffentlicht: (2024) -
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
von: Shu, Yan, et al.
Veröffentlicht: (2026) -
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
von: Zeng, Weichao, et al.
Veröffentlicht: (2024) -
TADoc: Robust Time-Aware Document Image Dewarping
von: Zhao, Fangmin, et al.
Veröffentlicht: (2025)