Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
Fuente:
arXiv
Guardado en:
| Autores principales: | Kišš, Martin, Hradiš, Michal |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Self-supervised Pre-training of Text Recognizers
por: Kišš, Martin, et al.
Publicado: (2024)
por: Kišš, Martin, et al.
Publicado: (2024)
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
por: Kišš, Martin, et al.
Publicado: (2025)
por: Kišš, Martin, et al.
Publicado: (2025)
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
por: Kohút, Jan, et al.
Publicado: (2025)
por: Kohút, Jan, et al.
Publicado: (2025)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
por: Cazenavette, George, et al.
Publicado: (2025)
por: Cazenavette, George, et al.
Publicado: (2025)
DiNO-Diffusion. Scaling Medical Diffusion via Self-Supervised Pre-Training
por: Jimenez-Perez, Guillermo, et al.
Publicado: (2024)
por: Jimenez-Perez, Guillermo, et al.
Publicado: (2024)
Scale Efficient Training for Large Datasets
por: Zhou, Qing, et al.
Publicado: (2025)
por: Zhou, Qing, et al.
Publicado: (2025)
Bridging Diversity and Uncertainty in Active learning with Self-Supervised Pre-Training
por: Doucet, Paul, et al.
Publicado: (2024)
por: Doucet, Paul, et al.
Publicado: (2024)
MaskOpt: A Large-Scale Mask Optimization Dataset to Advance AI in Integrated Circuit Manufacturing
por: Hu, Yuting, et al.
Publicado: (2025)
por: Hu, Yuting, et al.
Publicado: (2025)
Towards Writing Style Adaptation in Handwriting Recognition
por: Kohút, Jan, et al.
Publicado: (2023)
por: Kohút, Jan, et al.
Publicado: (2023)
Fast Training of Diffusion Models with Masked Transformers
por: Zheng, Hongkai, et al.
Publicado: (2023)
por: Zheng, Hongkai, et al.
Publicado: (2023)
Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning
por: Liu, Yuti, et al.
Publicado: (2024)
por: Liu, Yuti, et al.
Publicado: (2024)
Integration of Self-Supervised BYOL in Semi-Supervised Medical Image Recognition
por: Feng, Hao, et al.
Publicado: (2024)
por: Feng, Hao, et al.
Publicado: (2024)
Masked Generative Nested Transformers with Decode Time Scaling
por: Goyal, Sahil, et al.
Publicado: (2025)
por: Goyal, Sahil, et al.
Publicado: (2025)
Gradient-Sign Masking for Task Vector Transport Across Pre-Trained Models
por: Rinaldi, Filippo, et al.
Publicado: (2025)
por: Rinaldi, Filippo, et al.
Publicado: (2025)
Blockwise Self-Supervised Learning at Scale
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2023)
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2023)
Pre-training Vision Transformers with Formula-driven Supervised Learning
por: Kataoka, Hirokatsu, et al.
Publicado: (2022)
por: Kataoka, Hirokatsu, et al.
Publicado: (2022)
Semi-Supervised Masked Autoencoders: Unlocking Vision Transformer Potential with Limited Data
por: Faysal, Atik, et al.
Publicado: (2026)
por: Faysal, Atik, et al.
Publicado: (2026)
D$^3$epth: Self-Supervised Depth Estimation with Dynamic Mask in Dynamic Scenes
por: Chen, Siyu, et al.
Publicado: (2024)
por: Chen, Siyu, et al.
Publicado: (2024)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
por: Kang, Wonjun, et al.
Publicado: (2025)
por: Kang, Wonjun, et al.
Publicado: (2025)
DohaScript: A Large-Scale Multi-Writer Dataset for Continuous Handwritten Hindi Text
por: Singh, Kunwar Arpit, et al.
Publicado: (2026)
por: Singh, Kunwar Arpit, et al.
Publicado: (2026)
An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders
por: Lowe, Scott C., et al.
Publicado: (2024)
por: Lowe, Scott C., et al.
Publicado: (2024)
Masking Improves Contrastive Self-Supervised Learning for ConvNets, and Saliency Tells You Where
por: Chin, Zhi-Yi, et al.
Publicado: (2023)
por: Chin, Zhi-Yi, et al.
Publicado: (2023)
Intelligent Anomaly Detection for Lane Rendering Using Transformer with Self-Supervised Pre-Training and Customized Fine-Tuning
por: Dong, Yongqi, et al.
Publicado: (2023)
por: Dong, Yongqi, et al.
Publicado: (2023)
A Survey of the Self Supervised Learning Mechanisms for Vision Transformers
por: Khan, Asifullah, et al.
Publicado: (2024)
por: Khan, Asifullah, et al.
Publicado: (2024)
Look Through Masks: Towards Masked Face Recognition with De-Occlusion Distillation
por: Li, Chenyu, et al.
Publicado: (2024)
por: Li, Chenyu, et al.
Publicado: (2024)
Masked Face Recognition with Generative-to-Discriminative Representations
por: Ge, Shiming, et al.
Publicado: (2024)
por: Ge, Shiming, et al.
Publicado: (2024)
Beyond Labels: A Self-Supervised Framework with Masked Autoencoders and Random Cropping for Breast Cancer Subtype Classification
por: Chiocchetti, Annalisa, et al.
Publicado: (2024)
por: Chiocchetti, Annalisa, et al.
Publicado: (2024)
Erasing Self-Supervised Learning Backdoor by Cluster Activation Masking
por: Qian, Shengsheng, et al.
Publicado: (2023)
por: Qian, Shengsheng, et al.
Publicado: (2023)
NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields
por: Irshad, Muhammad Zubair, et al.
Publicado: (2024)
por: Irshad, Muhammad Zubair, et al.
Publicado: (2024)
EUDA: An Efficient Unsupervised Domain Adaptation via Self-Supervised Vision Transformer
por: Abedi, Ali, et al.
Publicado: (2024)
por: Abedi, Ali, et al.
Publicado: (2024)
OCT-SelfNet: A Self-Supervised Framework with Multi-Modal Datasets for Generalized and Robust Retinal Disease Detection
por: Jannat, Fatema-E, et al.
Publicado: (2024)
por: Jannat, Fatema-E, et al.
Publicado: (2024)
SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
por: Ye, Hanrong, et al.
Publicado: (2023)
por: Ye, Hanrong, et al.
Publicado: (2023)
Training-Only Heterogeneous Image-Patch-Text Graph Supervision for Advancing Few-Shot Learning Adapters
por: Mohammad, Mohammed Rahman Sherif Khan, et al.
Publicado: (2026)
por: Mohammad, Mohammed Rahman Sherif Khan, et al.
Publicado: (2026)
Kaputt: A Large-Scale Dataset for Visual Defect Detection
por: Höfer, Sebastian, et al.
Publicado: (2025)
por: Höfer, Sebastian, et al.
Publicado: (2025)
Soft Label Pruning and Quantization for Large-Scale Dataset Distillation
por: Lingao, Xiao, et al.
Publicado: (2026)
por: Lingao, Xiao, et al.
Publicado: (2026)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
por: Yu, Zhuoran, et al.
Publicado: (2023)
por: Yu, Zhuoran, et al.
Publicado: (2023)
Transformer-Based Self-Supervised Learning for Histopathological Classification of Ischemic Stroke Clot Origin
por: Yeh, K., et al.
Publicado: (2024)
por: Yeh, K., et al.
Publicado: (2024)
Robust Pre-Training of Medical Vision-and-Language Models with Domain-Invariant Multi-Modal Masked Reconstruction
por: Filvantorkaman, Melika, et al.
Publicado: (2026)
por: Filvantorkaman, Melika, et al.
Publicado: (2026)
DeiT-LT Distillation Strikes Back for Vision Transformer Training on Long-Tailed Datasets
por: Rangwani, Harsh, et al.
Publicado: (2024)
por: Rangwani, Harsh, et al.
Publicado: (2024)
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
por: Chen, Yi-Chia, et al.
Publicado: (2024)
por: Chen, Yi-Chia, et al.
Publicado: (2024)
Ejemplares similares
-
Self-supervised Pre-training of Text Recognizers
por: Kišš, Martin, et al.
Publicado: (2024) -
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
por: Kišš, Martin, et al.
Publicado: (2025) -
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
por: Kohút, Jan, et al.
Publicado: (2025) -
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
por: Cazenavette, George, et al.
Publicado: (2025) -
DiNO-Diffusion. Scaling Medical Diffusion via Self-Supervised Pre-Training
por: Jimenez-Perez, Guillermo, et al.
Publicado: (2024)