Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kišš, Martin, Hradiš, Michal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-supervised Pre-training of Text Recognizers
von: Kišš, Martin, et al.
Veröffentlicht: (2024)
von: Kišš, Martin, et al.
Veröffentlicht: (2024)
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
von: Kišš, Martin, et al.
Veröffentlicht: (2025)
von: Kišš, Martin, et al.
Veröffentlicht: (2025)
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
von: Kohút, Jan, et al.
Veröffentlicht: (2025)
von: Kohút, Jan, et al.
Veröffentlicht: (2025)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
von: Cazenavette, George, et al.
Veröffentlicht: (2025)
von: Cazenavette, George, et al.
Veröffentlicht: (2025)
DiNO-Diffusion. Scaling Medical Diffusion via Self-Supervised Pre-Training
von: Jimenez-Perez, Guillermo, et al.
Veröffentlicht: (2024)
von: Jimenez-Perez, Guillermo, et al.
Veröffentlicht: (2024)
Scale Efficient Training for Large Datasets
von: Zhou, Qing, et al.
Veröffentlicht: (2025)
von: Zhou, Qing, et al.
Veröffentlicht: (2025)
Bridging Diversity and Uncertainty in Active learning with Self-Supervised Pre-Training
von: Doucet, Paul, et al.
Veröffentlicht: (2024)
von: Doucet, Paul, et al.
Veröffentlicht: (2024)
MaskOpt: A Large-Scale Mask Optimization Dataset to Advance AI in Integrated Circuit Manufacturing
von: Hu, Yuting, et al.
Veröffentlicht: (2025)
von: Hu, Yuting, et al.
Veröffentlicht: (2025)
Towards Writing Style Adaptation in Handwriting Recognition
von: Kohút, Jan, et al.
Veröffentlicht: (2023)
von: Kohút, Jan, et al.
Veröffentlicht: (2023)
Fast Training of Diffusion Models with Masked Transformers
von: Zheng, Hongkai, et al.
Veröffentlicht: (2023)
von: Zheng, Hongkai, et al.
Veröffentlicht: (2023)
Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning
von: Liu, Yuti, et al.
Veröffentlicht: (2024)
von: Liu, Yuti, et al.
Veröffentlicht: (2024)
Integration of Self-Supervised BYOL in Semi-Supervised Medical Image Recognition
von: Feng, Hao, et al.
Veröffentlicht: (2024)
von: Feng, Hao, et al.
Veröffentlicht: (2024)
Masked Generative Nested Transformers with Decode Time Scaling
von: Goyal, Sahil, et al.
Veröffentlicht: (2025)
von: Goyal, Sahil, et al.
Veröffentlicht: (2025)
Gradient-Sign Masking for Task Vector Transport Across Pre-Trained Models
von: Rinaldi, Filippo, et al.
Veröffentlicht: (2025)
von: Rinaldi, Filippo, et al.
Veröffentlicht: (2025)
Blockwise Self-Supervised Learning at Scale
von: Siddiqui, Shoaib Ahmed, et al.
Veröffentlicht: (2023)
von: Siddiqui, Shoaib Ahmed, et al.
Veröffentlicht: (2023)
Pre-training Vision Transformers with Formula-driven Supervised Learning
von: Kataoka, Hirokatsu, et al.
Veröffentlicht: (2022)
von: Kataoka, Hirokatsu, et al.
Veröffentlicht: (2022)
Semi-Supervised Masked Autoencoders: Unlocking Vision Transformer Potential with Limited Data
von: Faysal, Atik, et al.
Veröffentlicht: (2026)
von: Faysal, Atik, et al.
Veröffentlicht: (2026)
D$^3$epth: Self-Supervised Depth Estimation with Dynamic Mask in Dynamic Scenes
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
von: Kang, Wonjun, et al.
Veröffentlicht: (2025)
von: Kang, Wonjun, et al.
Veröffentlicht: (2025)
DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi Text
von: Singh, Kunwar Arpit, et al.
Veröffentlicht: (2026)
von: Singh, Kunwar Arpit, et al.
Veröffentlicht: (2026)
An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders
von: Lowe, Scott C., et al.
Veröffentlicht: (2024)
von: Lowe, Scott C., et al.
Veröffentlicht: (2024)
Masking Improves Contrastive Self-Supervised Learning for ConvNets, and Saliency Tells You Where
von: Chin, Zhi-Yi, et al.
Veröffentlicht: (2023)
von: Chin, Zhi-Yi, et al.
Veröffentlicht: (2023)
Intelligent Anomaly Detection for Lane Rendering Using Transformer with Self-Supervised Pre-Training and Customized Fine-Tuning
von: Dong, Yongqi, et al.
Veröffentlicht: (2023)
von: Dong, Yongqi, et al.
Veröffentlicht: (2023)
A Survey of the Self Supervised Learning Mechanisms for Vision Transformers
von: Khan, Asifullah, et al.
Veröffentlicht: (2024)
von: Khan, Asifullah, et al.
Veröffentlicht: (2024)
Look Through Masks: Towards Masked Face Recognition with De-Occlusion Distillation
von: Li, Chenyu, et al.
Veröffentlicht: (2024)
von: Li, Chenyu, et al.
Veröffentlicht: (2024)
Masked Face Recognition with Generative-to-Discriminative Representations
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
Beyond Labels: A Self-Supervised Framework with Masked Autoencoders and Random Cropping for Breast Cancer Subtype Classification
von: Chiocchetti, Annalisa, et al.
Veröffentlicht: (2024)
von: Chiocchetti, Annalisa, et al.
Veröffentlicht: (2024)
Erasing Self-Supervised Learning Backdoor by Cluster Activation Masking
von: Qian, Shengsheng, et al.
Veröffentlicht: (2023)
von: Qian, Shengsheng, et al.
Veröffentlicht: (2023)
NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields
von: Irshad, Muhammad Zubair, et al.
Veröffentlicht: (2024)
von: Irshad, Muhammad Zubair, et al.
Veröffentlicht: (2024)
EUDA: An Efficient Unsupervised Domain Adaptation via Self-Supervised Vision Transformer
von: Abedi, Ali, et al.
Veröffentlicht: (2024)
von: Abedi, Ali, et al.
Veröffentlicht: (2024)
OCT-SelfNet: A Self-Supervised Framework with Multi-Modal Datasets for Generalized and Robust Retinal Disease Detection
von: Jannat, Fatema-E, et al.
Veröffentlicht: (2024)
von: Jannat, Fatema-E, et al.
Veröffentlicht: (2024)
SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
von: Ye, Hanrong, et al.
Veröffentlicht: (2023)
von: Ye, Hanrong, et al.
Veröffentlicht: (2023)
Training-Only Heterogeneous Image-Patch-Text Graph Supervision for Advancing Few-Shot Learning Adapters
von: Mohammad, Mohammed Rahman Sherif Khan, et al.
Veröffentlicht: (2026)
von: Mohammad, Mohammed Rahman Sherif Khan, et al.
Veröffentlicht: (2026)
Kaputt: A Large-Scale Dataset for Visual Defect Detection
von: Höfer, Sebastian, et al.
Veröffentlicht: (2025)
von: Höfer, Sebastian, et al.
Veröffentlicht: (2025)
Soft Label Pruning and Quantization for Large-Scale Dataset Distillation
von: Lingao, Xiao, et al.
Veröffentlicht: (2026)
von: Lingao, Xiao, et al.
Veröffentlicht: (2026)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
von: Yu, Zhuoran, et al.
Veröffentlicht: (2023)
von: Yu, Zhuoran, et al.
Veröffentlicht: (2023)
Transformer-Based Self-Supervised Learning for Histopathological Classification of Ischemic Stroke Clot Origin
von: Yeh, K., et al.
Veröffentlicht: (2024)
von: Yeh, K., et al.
Veröffentlicht: (2024)
Robust Pre-Training of Medical Vision-and-Language Models with Domain-Invariant Multi-Modal Masked Reconstruction
von: Filvantorkaman, Melika, et al.
Veröffentlicht: (2026)
von: Filvantorkaman, Melika, et al.
Veröffentlicht: (2026)
DeiT-LT Distillation Strikes Back for Vision Transformer Training on Long-Tailed Datasets
von: Rangwani, Harsh, et al.
Veröffentlicht: (2024)
von: Rangwani, Harsh, et al.
Veröffentlicht: (2024)
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Self-supervised Pre-training of Text Recognizers
von: Kišš, Martin, et al.
Veröffentlicht: (2024) -
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
von: Kišš, Martin, et al.
Veröffentlicht: (2025) -
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
von: Kohút, Jan, et al.
Veröffentlicht: (2025) -
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
von: Cazenavette, George, et al.
Veröffentlicht: (2025) -
DiNO-Diffusion. Scaling Medical Diffusion via Self-Supervised Pre-Training
von: Jimenez-Perez, Guillermo, et al.
Veröffentlicht: (2024)