Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Zuan, Wang, Yuxin, Qu, Yadong, Zhang, Boqiang, Wang, Zixiao, Xu, Jianjun, Xie, Hongtao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing
by: Zhang, Boqiang, et al.
Published: (2024)
by: Zhang, Boqiang, et al.
Published: (2024)
Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition
by: Zhou, Bangbang, et al.
Published: (2024)
by: Zhou, Bangbang, et al.
Published: (2024)
How Control Information Influences Multilingual Text Image Generation and Editing?
by: Zhang, Boqiang, et al.
Published: (2024)
by: Zhang, Boqiang, et al.
Published: (2024)
Boosting Semi-Supervised Scene Text Recognition via Viewing and Summarizing
by: Qu, Yadong, et al.
Published: (2024)
by: Qu, Yadong, et al.
Published: (2024)
Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement
by: Guo, Junrong, et al.
Published: (2026)
by: Guo, Junrong, et al.
Published: (2026)
SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition
by: Du, Yongkun, et al.
Published: (2024)
by: Du, Yongkun, et al.
Published: (2024)
IGD: Instructional Graphic Design with Multimodal Layer Generation
by: Qu, Yadong, et al.
Published: (2025)
by: Qu, Yadong, et al.
Published: (2025)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition
by: Haliassos, Alexandros, et al.
Published: (2024)
by: Haliassos, Alexandros, et al.
Published: (2024)
Bridging Synthetic and Real Worlds for Pre-training Scene Text Detectors
by: Guan, Tongkun, et al.
Published: (2023)
by: Guan, Tongkun, et al.
Published: (2023)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Decoder Pre-Training with only Text for Scene Text Recognition
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Self-Supervised Pre-Training for Table Structure Recognition Transformer
by: Peng, ShengYun, et al.
Published: (2024)
by: Peng, ShengYun, et al.
Published: (2024)
The First Swahili Language Scene Text Detection and Recognition Dataset
by: Douamba, Fadila Wendigoundi, et al.
Published: (2024)
by: Douamba, Fadila Wendigoundi, et al.
Published: (2024)
In Pursuit of Pixel Supervision for Visual Pre-training
by: Yang, Lihe, et al.
Published: (2025)
by: Yang, Lihe, et al.
Published: (2025)
Enhancing SAR Object Detection with Self-Supervised Pre-training on Masked Auto-Encoders
by: Pu, Xinyang, et al.
Published: (2025)
by: Pu, Xinyang, et al.
Published: (2025)
Self-Supervised Pre-training with Combined Datasets for 3D Perception in Autonomous Driving
by: Wang, Shumin, et al.
Published: (2025)
by: Wang, Shumin, et al.
Published: (2025)
Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
by: Ebouky, Brown, et al.
Published: (2025)
by: Ebouky, Brown, et al.
Published: (2025)
LEGO: Self-Supervised Representation Learning for Scene Text Images
by: Ren, Yujin, et al.
Published: (2024)
by: Ren, Yujin, et al.
Published: (2024)
From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition
by: Cai, Chen, et al.
Published: (2025)
by: Cai, Chen, et al.
Published: (2025)
Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
by: Zhang, Yifei, et al.
Published: (2025)
by: Zhang, Yifei, et al.
Published: (2025)
ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting
by: Duan, Chen, et al.
Published: (2024)
by: Duan, Chen, et al.
Published: (2024)
Effect of Rotation Angle in Self-Supervised Pre-training is Dataset-Dependent
by: Saranchuk, Amy, et al.
Published: (2024)
by: Saranchuk, Amy, et al.
Published: (2024)
Endo-CLIP: Progressive Self-Supervised Pre-training on Raw Colonoscopy Records
by: He, Yili, et al.
Published: (2025)
by: He, Yili, et al.
Published: (2025)
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
by: Lei, Jiachen, et al.
Published: (2025)
by: Lei, Jiachen, et al.
Published: (2025)
N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
Enhancing Vision-Language Pre-training with Rich Supervisions
by: Gao, Yuan, et al.
Published: (2024)
by: Gao, Yuan, et al.
Published: (2024)
Self-Supervised Learning for Text Recognition: A Critical Survey
by: Penarrubia, Carlos, et al.
Published: (2024)
by: Penarrubia, Carlos, et al.
Published: (2024)
Micro-Expression Recognition by Motion Feature Extraction based on Pre-training
by: Li, Ruolin, et al.
Published: (2024)
by: Li, Ruolin, et al.
Published: (2024)
PreGSU-A Generalized Traffic Scene Understanding Model for Autonomous Driving based on Pre-trained Graph Attention Network
by: Wang, Yuning, et al.
Published: (2024)
by: Wang, Yuning, et al.
Published: (2024)
PatchContrast: Self-Supervised Pre-training for 3D Object Detection
by: Shrout, Oren, et al.
Published: (2023)
by: Shrout, Oren, et al.
Published: (2023)
A Closer Look at Benchmarking Self-Supervised Pre-training with Image Classification
by: Marks, Markus, et al.
Published: (2024)
by: Marks, Markus, et al.
Published: (2024)
TextMastero: Mastering High-Quality Scene Text Editing in Diverse Languages and Styles
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
Weakly Supervised Lymph Nodes Segmentation Based on Partial Instance Annotations with Pre-trained Dual-branch Network and Pseudo Label Learning
by: Wang, Litingyu, et al.
Published: (2024)
by: Wang, Litingyu, et al.
Published: (2024)
Masked Next-Scale Prediction for Self-supervised Scene Text Recognition
by: Chen, Zhuohao, et al.
Published: (2026)
by: Chen, Zhuohao, et al.
Published: (2026)
Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition
by: Lin, Tiancheng, et al.
Published: (2024)
by: Lin, Tiancheng, et al.
Published: (2024)
Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors
by: Xu, Peiran, et al.
Published: (2025)
by: Xu, Peiran, et al.
Published: (2025)
Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis
by: Bi, Tianci, et al.
Published: (2024)
by: Bi, Tianci, et al.
Published: (2024)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
by: Ma, Shuailei, et al.
Published: (2023)
by: Ma, Shuailei, et al.
Published: (2023)
Similar Items
-
Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing
by: Zhang, Boqiang, et al.
Published: (2024) -
Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition
by: Zhou, Bangbang, et al.
Published: (2024) -
How Control Information Influences Multilingual Text Image Generation and Editing?
by: Zhang, Boqiang, et al.
Published: (2024) -
Boosting Semi-Supervised Scene Text Recognition via Viewing and Summarizing
by: Qu, Yadong, et al.
Published: (2024) -
Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling
by: Wang, Zixiao, et al.
Published: (2024)