Self-supervised Pre-training of Text Recognizers
Fuente:
arXiv
Salvato in:
| Autori principali: | Kišš, Martin, Hradiš, Michal |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
di: Kišš, Martin, et al.
Pubblicazione: (2025)
di: Kišš, Martin, et al.
Pubblicazione: (2025)
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
di: Kišš, Martin, et al.
Pubblicazione: (2025)
di: Kišš, Martin, et al.
Pubblicazione: (2025)
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
di: Kohút, Jan, et al.
Pubblicazione: (2025)
di: Kohút, Jan, et al.
Pubblicazione: (2025)
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
di: Bawazir, Ameera, et al.
Pubblicazione: (2024)
di: Bawazir, Ameera, et al.
Pubblicazione: (2024)
SST: Self-training with Self-adaptive Thresholding for Semi-supervised Learning
di: Zhao, Shuai, et al.
Pubblicazione: (2025)
di: Zhao, Shuai, et al.
Pubblicazione: (2025)
Inverse-LLaVA: Eliminating Alignment Pre-training Through Text-to-Vision Mapping
di: Zhan, Xuhui, et al.
Pubblicazione: (2025)
di: Zhan, Xuhui, et al.
Pubblicazione: (2025)
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
di: Rao, Zhefan, et al.
Pubblicazione: (2024)
di: Rao, Zhefan, et al.
Pubblicazione: (2024)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
di: Kim, Sanghwan, et al.
Pubblicazione: (2024)
di: Kim, Sanghwan, et al.
Pubblicazione: (2024)
LightFair: Towards an Efficient Alternative for Fair T2I Diffusion via Debiasing Pre-trained Text Encoders
di: Han, Boyu, et al.
Pubblicazione: (2025)
di: Han, Boyu, et al.
Pubblicazione: (2025)
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
di: Zhang, Yuhui, et al.
Pubblicazione: (2023)
di: Zhang, Yuhui, et al.
Pubblicazione: (2023)
Feature Hallucination for Self-supervised Action Recognition
di: Wang, Lei, et al.
Pubblicazione: (2025)
di: Wang, Lei, et al.
Pubblicazione: (2025)
Self-supervised Learning for Hyperspectral Images of Trees
di: Rahman, Moqsadur, et al.
Pubblicazione: (2025)
di: Rahman, Moqsadur, et al.
Pubblicazione: (2025)
Self-supervised Transformation Learning for Equivariant Representations
di: Yu, Jaemyung, et al.
Pubblicazione: (2025)
di: Yu, Jaemyung, et al.
Pubblicazione: (2025)
Dataset Ownership Verification in Contrastive Pre-trained Models
di: Xie, Yuechen, et al.
Pubblicazione: (2025)
di: Xie, Yuechen, et al.
Pubblicazione: (2025)
Practical Continual Forgetting for Pre-trained Vision Models
di: Zhao, Hongbo, et al.
Pubblicazione: (2025)
di: Zhao, Hongbo, et al.
Pubblicazione: (2025)
Pre-training with Random Orthogonal Projection Image Modeling
di: Haghighat, Maryam, et al.
Pubblicazione: (2023)
di: Haghighat, Maryam, et al.
Pubblicazione: (2023)
Multi-modal Vision Pre-training for Medical Image Analysis
di: Rui, Shaohao, et al.
Pubblicazione: (2024)
di: Rui, Shaohao, et al.
Pubblicazione: (2024)
Vehicle-centric Perception via Multimodal Structured Pre-training
di: Wu, Wentao, et al.
Pubblicazione: (2025)
di: Wu, Wentao, et al.
Pubblicazione: (2025)
Stylized Structural Patterns for Improved Neural Network Pre-training
di: Salehi, Farnood, et al.
Pubblicazione: (2025)
di: Salehi, Farnood, et al.
Pubblicazione: (2025)
Pre-training Vision Transformers with Formula-driven Supervised Learning
di: Kataoka, Hirokatsu, et al.
Pubblicazione: (2022)
di: Kataoka, Hirokatsu, et al.
Pubblicazione: (2022)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
di: Berrada, Tariq, et al.
Pubblicazione: (2023)
di: Berrada, Tariq, et al.
Pubblicazione: (2023)
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
di: Chen, Hao, et al.
Pubblicazione: (2023)
di: Chen, Hao, et al.
Pubblicazione: (2023)
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
di: Gupta, Gunshi, et al.
Pubblicazione: (2024)
di: Gupta, Gunshi, et al.
Pubblicazione: (2024)
Towards Writing Style Adaptation in Handwriting Recognition
di: Kohút, Jan, et al.
Pubblicazione: (2023)
di: Kohút, Jan, et al.
Pubblicazione: (2023)
Slight Corruption in Pre-training Data Makes Better Diffusion Models
di: Chen, Hao, et al.
Pubblicazione: (2024)
di: Chen, Hao, et al.
Pubblicazione: (2024)
Gradient-based Fine-Tuning through Pre-trained Model Regularization
di: Liu, Xuanbo, et al.
Pubblicazione: (2025)
di: Liu, Xuanbo, et al.
Pubblicazione: (2025)
Benchmarking the Influence of Pre-training on Explanation Performance in MR Image Classification
di: Oliveira, Marta, et al.
Pubblicazione: (2023)
di: Oliveira, Marta, et al.
Pubblicazione: (2023)
GPD-1: Generative Pre-training for Driving
di: Xie, Zixun, et al.
Pubblicazione: (2024)
di: Xie, Zixun, et al.
Pubblicazione: (2024)
SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
di: Zhang, Gengwei, et al.
Pubblicazione: (2024)
di: Zhang, Gengwei, et al.
Pubblicazione: (2024)
Domain-Specific Pre-training Improves Confidence in Whole Slide Image Classification
di: Chitnis, Soham Rohit, et al.
Pubblicazione: (2023)
di: Chitnis, Soham Rohit, et al.
Pubblicazione: (2023)
Effective Backdoor Mitigation in Vision-Language Models Depends on the Pre-training Objective
di: Verma, Sahil, et al.
Pubblicazione: (2023)
di: Verma, Sahil, et al.
Pubblicazione: (2023)
Region-Aware Reconstruction Strategy for Pre-training fMRI Foundation Model
di: Doodipala, Ruthwik Reddy, et al.
Pubblicazione: (2025)
di: Doodipala, Ruthwik Reddy, et al.
Pubblicazione: (2025)
Closed-Form Linear-Probe Dataset Distillation for Pre-trained Vision Models
di: Peng, Bincheng, et al.
Pubblicazione: (2026)
di: Peng, Bincheng, et al.
Pubblicazione: (2026)
Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision
di: Dong, Hao, et al.
Pubblicazione: (2024)
di: Dong, Hao, et al.
Pubblicazione: (2024)
Federated Learning for Face Recognition via Intra-subject Self-supervised Learning
di: Kim, Hansol, et al.
Pubblicazione: (2024)
di: Kim, Hansol, et al.
Pubblicazione: (2024)
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
di: Gidaris, Spyros, et al.
Pubblicazione: (2023)
di: Gidaris, Spyros, et al.
Pubblicazione: (2023)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
di: Lai, Zhengfeng, et al.
Pubblicazione: (2024)
di: Lai, Zhengfeng, et al.
Pubblicazione: (2024)
Multi-View and Multi-Scale Alignment for Contrastive Language-Image Pre-training in Mammography
di: Du, Yuexi, et al.
Pubblicazione: (2024)
di: Du, Yuexi, et al.
Pubblicazione: (2024)
Feedback-based Modal Mutual Search for Attacking Vision-Language Pre-training Models
di: Ding, Renhua, et al.
Pubblicazione: (2024)
di: Ding, Renhua, et al.
Pubblicazione: (2024)
Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models
di: Tang, Yiwen, et al.
Pubblicazione: (2023)
di: Tang, Yiwen, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Masked Self-Supervised Pre-Training for Text Recognition Transformers on Large-Scale Datasets
di: Kišš, Martin, et al.
Pubblicazione: (2025) -
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
di: Kišš, Martin, et al.
Pubblicazione: (2025) -
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction
di: Kohút, Jan, et al.
Pubblicazione: (2025) -
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
di: Bawazir, Ameera, et al.
Pubblicazione: (2024) -
SST: Self-training with Self-adaptive Thresholding for Semi-supervised Learning
di: Zhao, Shuai, et al.
Pubblicazione: (2025)