Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Kišš, Martin, Hradiš, Michal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-supervised Pre-training of Text Recognizers
by: Kišš, Martin, et al.
Published: (2024)
by: Kišš, Martin, et al.
Published: (2024)
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
by: Kišš, Martin, et al.
Published: (2025)
by: Kišš, Martin, et al.
Published: (2025)
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
by: Kohút, Jan, et al.
Published: (2025)
by: Kohút, Jan, et al.
Published: (2025)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
DiNO-Diffusion. Scaling Medical Diffusion via Self-Supervised Pre-Training
by: Jimenez-Perez, Guillermo, et al.
Published: (2024)
by: Jimenez-Perez, Guillermo, et al.
Published: (2024)
Scale Efficient Training for Large Datasets
by: Zhou, Qing, et al.
Published: (2025)
by: Zhou, Qing, et al.
Published: (2025)
Bridging Diversity and Uncertainty in Active learning with Self-Supervised Pre-Training
by: Doucet, Paul, et al.
Published: (2024)
by: Doucet, Paul, et al.
Published: (2024)
MaskOpt: A Large-Scale Mask Optimization Dataset to Advance AI in Integrated Circuit Manufacturing
by: Hu, Yuting, et al.
Published: (2025)
by: Hu, Yuting, et al.
Published: (2025)
Towards Writing Style Adaptation in Handwriting Recognition
by: Kohút, Jan, et al.
Published: (2023)
by: Kohút, Jan, et al.
Published: (2023)
Fast Training of Diffusion Models with Masked Transformers
by: Zheng, Hongkai, et al.
Published: (2023)
by: Zheng, Hongkai, et al.
Published: (2023)
Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning
by: Liu, Yuti, et al.
Published: (2024)
by: Liu, Yuti, et al.
Published: (2024)
Integration of Self-Supervised BYOL in Semi-Supervised Medical Image Recognition
by: Feng, Hao, et al.
Published: (2024)
by: Feng, Hao, et al.
Published: (2024)
Masked Generative Nested Transformers with Decode Time Scaling
by: Goyal, Sahil, et al.
Published: (2025)
by: Goyal, Sahil, et al.
Published: (2025)
Gradient-Sign Masking for Task Vector Transport Across Pre-Trained Models
by: Rinaldi, Filippo, et al.
Published: (2025)
by: Rinaldi, Filippo, et al.
Published: (2025)
Blockwise Self-Supervised Learning at Scale
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
Pre-training Vision Transformers with Formula-driven Supervised Learning
by: Kataoka, Hirokatsu, et al.
Published: (2022)
by: Kataoka, Hirokatsu, et al.
Published: (2022)
Semi-Supervised Masked Autoencoders: Unlocking Vision Transformer Potential with Limited Data
by: Faysal, Atik, et al.
Published: (2026)
by: Faysal, Atik, et al.
Published: (2026)
D$^3$epth: Self-Supervised Depth Estimation with Dynamic Mask in Dynamic Scenes
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi Text
by: Singh, Kunwar Arpit, et al.
Published: (2026)
by: Singh, Kunwar Arpit, et al.
Published: (2026)
An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders
by: Lowe, Scott C., et al.
Published: (2024)
by: Lowe, Scott C., et al.
Published: (2024)
Masking Improves Contrastive Self-Supervised Learning for ConvNets, and Saliency Tells You Where
by: Chin, Zhi-Yi, et al.
Published: (2023)
by: Chin, Zhi-Yi, et al.
Published: (2023)
Intelligent Anomaly Detection for Lane Rendering Using Transformer with Self-Supervised Pre-Training and Customized Fine-Tuning
by: Dong, Yongqi, et al.
Published: (2023)
by: Dong, Yongqi, et al.
Published: (2023)
A Survey of the Self Supervised Learning Mechanisms for Vision Transformers
by: Khan, Asifullah, et al.
Published: (2024)
by: Khan, Asifullah, et al.
Published: (2024)
Look Through Masks: Towards Masked Face Recognition with De-Occlusion Distillation
by: Li, Chenyu, et al.
Published: (2024)
by: Li, Chenyu, et al.
Published: (2024)
Masked Face Recognition with Generative-to-Discriminative Representations
by: Ge, Shiming, et al.
Published: (2024)
by: Ge, Shiming, et al.
Published: (2024)
Beyond Labels: A Self-Supervised Framework with Masked Autoencoders and Random Cropping for Breast Cancer Subtype Classification
by: Chiocchetti, Annalisa, et al.
Published: (2024)
by: Chiocchetti, Annalisa, et al.
Published: (2024)
Erasing Self-Supervised Learning Backdoor by Cluster Activation Masking
by: Qian, Shengsheng, et al.
Published: (2023)
by: Qian, Shengsheng, et al.
Published: (2023)
NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields
by: Irshad, Muhammad Zubair, et al.
Published: (2024)
by: Irshad, Muhammad Zubair, et al.
Published: (2024)
EUDA: An Efficient Unsupervised Domain Adaptation via Self-Supervised Vision Transformer
by: Abedi, Ali, et al.
Published: (2024)
by: Abedi, Ali, et al.
Published: (2024)
OCT-SelfNet: A Self-Supervised Framework with Multi-Modal Datasets for Generalized and Robust Retinal Disease Detection
by: Jannat, Fatema-E, et al.
Published: (2024)
by: Jannat, Fatema-E, et al.
Published: (2024)
SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
by: Ye, Hanrong, et al.
Published: (2023)
by: Ye, Hanrong, et al.
Published: (2023)
Training-Only Heterogeneous Image-Patch-Text Graph Supervision for Advancing Few-Shot Learning Adapters
by: Mohammad, Mohammed Rahman Sherif Khan, et al.
Published: (2026)
by: Mohammad, Mohammed Rahman Sherif Khan, et al.
Published: (2026)
Kaputt: A Large-Scale Dataset for Visual Defect Detection
by: Höfer, Sebastian, et al.
Published: (2025)
by: Höfer, Sebastian, et al.
Published: (2025)
Soft Label Pruning and Quantization for Large-Scale Dataset Distillation
by: Lingao, Xiao, et al.
Published: (2026)
by: Lingao, Xiao, et al.
Published: (2026)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
by: Yu, Zhuoran, et al.
Published: (2023)
by: Yu, Zhuoran, et al.
Published: (2023)
Transformer-Based Self-Supervised Learning for Histopathological Classification of Ischemic Stroke Clot Origin
by: Yeh, K., et al.
Published: (2024)
by: Yeh, K., et al.
Published: (2024)
Robust Pre-Training of Medical Vision-and-Language Models with Domain-Invariant Multi-Modal Masked Reconstruction
by: Filvantorkaman, Melika, et al.
Published: (2026)
by: Filvantorkaman, Melika, et al.
Published: (2026)
DeiT-LT Distillation Strikes Back for Vision Transformer Training on Long-Tailed Datasets
by: Rangwani, Harsh, et al.
Published: (2024)
by: Rangwani, Harsh, et al.
Published: (2024)
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
by: Chen, Yi-Chia, et al.
Published: (2024)
by: Chen, Yi-Chia, et al.
Published: (2024)
Similar Items
-
Self-supervised Pre-training of Text Recognizers
by: Kišš, Martin, et al.
Published: (2024) -
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
by: Kišš, Martin, et al.
Published: (2025) -
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
by: Kohút, Jan, et al.
Published: (2025) -
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025) -
DiNO-Diffusion. Scaling Medical Diffusion via Self-Supervised Pre-Training
by: Jimenez-Perez, Guillermo, et al.
Published: (2024)