LEGO: Self-Supervised Representation Learning for Scene Text Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Yujin, Zhang, Jiaxin, Jin, Lianwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Online Writer Retrieval with Chinese Handwritten Phrases: A Synergistic Temporal-Frequency Representation Learning Approach
von: Zhang, Peirong, et al.
Veröffentlicht: (2024)
von: Zhang, Peirong, et al.
Veröffentlicht: (2024)
Revisiting Tampered Scene Text Detection in the Era of Generative AI
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
von: Peng, Dezhi, et al.
Veröffentlicht: (2023)
von: Peng, Dezhi, et al.
Veröffentlicht: (2023)
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
Progressive Evolution from Single-Point to Polygon for Scene Text
von: Deng, Linger, et al.
Veröffentlicht: (2023)
von: Deng, Linger, et al.
Veröffentlicht: (2023)
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
Capturing More: Learning Multi-Domain Representations for Robust Online Handwriting Verification
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
von: Qu, Chenfan, et al.
Veröffentlicht: (2025)
von: Qu, Chenfan, et al.
Veröffentlicht: (2025)
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
Self-Supervised Learning of Plant Image Representations
von: Moummad, Ilyass, et al.
Veröffentlicht: (2026)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2026)
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
Open-Vocabulary Scene Text Recognition via Pseudo-Image Labeling and Margin Loss
von: Ren, Xuhua, et al.
Veröffentlicht: (2024)
von: Ren, Xuhua, et al.
Veröffentlicht: (2024)
Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
von: Zhang, Yifei, et al.
Veröffentlicht: (2025)
von: Zhang, Yifei, et al.
Veröffentlicht: (2025)
Self-Supervised Scene Flow Estimation with Point-Voxel Fusion and Surface Representation
von: Xiang, Xuezhi, et al.
Veröffentlicht: (2024)
von: Xiang, Xuezhi, et al.
Veröffentlicht: (2024)
Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
von: Gao, Zuan, et al.
Veröffentlicht: (2024)
von: Gao, Zuan, et al.
Veröffentlicht: (2024)
In-Domain Self-Supervised Learning Improves Remote Sensing Image Scene Classification
von: Dimitrovski, Ivica, et al.
Veröffentlicht: (2023)
von: Dimitrovski, Ivica, et al.
Veröffentlicht: (2023)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
von: Watawana, Hasindri, et al.
Veröffentlicht: (2024)
von: Watawana, Hasindri, et al.
Veröffentlicht: (2024)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Self-Supervised High Dynamic Range Imaging with Multi-Exposure Images in Dynamic Scenes
von: Zhang, Zhilu, et al.
Veröffentlicht: (2023)
von: Zhang, Zhilu, et al.
Veröffentlicht: (2023)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
von: Li, Hongliang, et al.
Veröffentlicht: (2025)
von: Li, Hongliang, et al.
Veröffentlicht: (2025)
MCCD: A Multi-Attribute Chinese Calligraphy Character Dataset Annotated with Script Styles, Dynasties, and Calligraphers
von: Zhao, Yixin, et al.
Veröffentlicht: (2025)
von: Zhao, Yixin, et al.
Veröffentlicht: (2025)
TextShield-R1: Reinforced Reasoning for Tampered Text Detection
von: Qu, Chenfan, et al.
Veröffentlicht: (2026)
von: Qu, Chenfan, et al.
Veröffentlicht: (2026)
On the Discriminability of Self-Supervised Representation Learning
von: Song, Zeen, et al.
Veröffentlicht: (2024)
von: Song, Zeen, et al.
Veröffentlicht: (2024)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
Equivariant Representation Learning for Augmentation-based Self-Supervised Learning via Image Reconstruction
von: Wang, Qin, et al.
Veröffentlicht: (2024)
von: Wang, Qin, et al.
Veröffentlicht: (2024)
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning
von: Chen, Yiyang, et al.
Veröffentlicht: (2025)
von: Chen, Yiyang, et al.
Veröffentlicht: (2025)
Information-Maximized Soft Variable Discretization for Self-Supervised Image Representation Learning
von: Niu, Chuang, et al.
Veröffentlicht: (2025)
von: Niu, Chuang, et al.
Veröffentlicht: (2025)
TextSleuth: Towards Explainable Tampered Text Detection
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
von: Pan, Wei, et al.
Veröffentlicht: (2025)
von: Pan, Wei, et al.
Veröffentlicht: (2025)
Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision
von: Gao, Yunhe, et al.
Veröffentlicht: (2026)
von: Gao, Yunhe, et al.
Veröffentlicht: (2026)
Self-Supervised Learning Based on Transformed Image Reconstruction for Equivariance-Coherent Feature Representation
von: Wang, Qin, et al.
Veröffentlicht: (2025)
von: Wang, Qin, et al.
Veröffentlicht: (2025)
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
von: Ahamed, Shihab Aaqil, et al.
Veröffentlicht: (2025)
von: Ahamed, Shihab Aaqil, et al.
Veröffentlicht: (2025)
Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition
von: Lin, Tiancheng, et al.
Veröffentlicht: (2024)
von: Lin, Tiancheng, et al.
Veröffentlicht: (2024)
Self-Supervised Representation Learning for Adversarial Attack Detection
von: Li, Yi, et al.
Veröffentlicht: (2024)
von: Li, Yi, et al.
Veröffentlicht: (2024)
Self-Supervised Representation Learning with Meta Comprehensive Regularization
von: Guo, Huijie, et al.
Veröffentlicht: (2024)
von: Guo, Huijie, et al.
Veröffentlicht: (2024)
Sonata: Self-Supervised Learning of Reliable Point Representations
von: Wu, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyang, et al.
Veröffentlicht: (2025)
Enhancing Representations through Heterogeneous Self-Supervised Learning
von: Li, Zhong-Yu, et al.
Veröffentlicht: (2023)
von: Li, Zhong-Yu, et al.
Veröffentlicht: (2023)
Boosting Semi-Supervised Scene Text Recognition via Viewing and Summarizing
von: Qu, Yadong, et al.
Veröffentlicht: (2024)
von: Qu, Yadong, et al.
Veröffentlicht: (2024)
AFiRe: Anatomy-Driven Self-Supervised Learning for Fine-Grained Representation in Radiographic Images
von: Liu, Yihang, et al.
Veröffentlicht: (2025)
von: Liu, Yihang, et al.
Veröffentlicht: (2025)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Online Writer Retrieval with Chinese Handwritten Phrases: A Synergistic Temporal-Frequency Representation Learning Approach
von: Zhang, Peirong, et al.
Veröffentlicht: (2024) -
Revisiting Tampered Scene Text Detection in the Era of Generative AI
von: Qu, Chenfan, et al.
Veröffentlicht: (2024) -
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
von: Peng, Dezhi, et al.
Veröffentlicht: (2023) -
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024) -
Progressive Evolution from Single-Point to Polygon for Scene Text
von: Deng, Linger, et al.
Veröffentlicht: (2023)