CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Shuai, Quan, Ruijie, Zhu, Linchao, Yang, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
MDiff4STR: Mask Diffusion Model for Scene Text Recognition
von: Du, Yongkun, et al.
Veröffentlicht: (2025)
von: Du, Yongkun, et al.
Veröffentlicht: (2025)
AudioScenic: Audio-Driven Video Scene Editing
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
von: Feng, Tuo, et al.
Veröffentlicht: (2024)
von: Feng, Tuo, et al.
Veröffentlicht: (2024)
Decoder Pre-Training with only Text for Scene Text Recognition
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
IndicSTR12: A Dataset for Indic Scene Text Recognition
von: Lunia, Harsh, et al.
Veröffentlicht: (2024)
von: Lunia, Harsh, et al.
Veröffentlicht: (2024)
MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training
von: Zhu, Lei, et al.
Veröffentlicht: (2025)
von: Zhu, Lei, et al.
Veröffentlicht: (2025)
DiffSTR: Controlled Diffusion Models for Scene Text Removal
von: Pathak, Sanhita, et al.
Veröffentlicht: (2024)
von: Pathak, Sanhita, et al.
Veröffentlicht: (2024)
Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
von: Gao, Zuan, et al.
Veröffentlicht: (2024)
von: Gao, Zuan, et al.
Veröffentlicht: (2024)
Slimmable Networks for Contrastive Self-supervised Learning
von: Zhao, Shuai, et al.
Veröffentlicht: (2022)
von: Zhao, Shuai, et al.
Veröffentlicht: (2022)
ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
von: Yao, Xin, et al.
Veröffentlicht: (2025)
von: Yao, Xin, et al.
Veröffentlicht: (2025)
FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
von: Lan, Rui, et al.
Veröffentlicht: (2025)
von: Lan, Rui, et al.
Veröffentlicht: (2025)
Endo-CLIP: Progressive Self-Supervised Pre-training on Raw Colonoscopy Records
von: He, Yili, et al.
Veröffentlicht: (2025)
von: He, Yili, et al.
Veröffentlicht: (2025)
DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)
3D Scene Graph Guided Vision-Language Pre-training
von: Liu, Hao, et al.
Veröffentlicht: (2024)
von: Liu, Hao, et al.
Veröffentlicht: (2024)
Bridging Synthetic and Real Worlds for Pre-training Scene Text Detectors
von: Guan, Tongkun, et al.
Veröffentlicht: (2023)
von: Guan, Tongkun, et al.
Veröffentlicht: (2023)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
STR-Cert: Robustness Certification for Deep Text Recognition on Deep Learning Pipelines and Vision Transformers
von: Shao, Daqian, et al.
Veröffentlicht: (2023)
von: Shao, Daqian, et al.
Veröffentlicht: (2023)
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
GaitSTR: Gait Recognition with Sequential Two-stream Refinement
von: Zheng, Wanrong, et al.
Veröffentlicht: (2024)
von: Zheng, Wanrong, et al.
Veröffentlicht: (2024)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
Continual Forgetting for Pre-trained Vision Models
von: Zhao, Hongbo, et al.
Veröffentlicht: (2024)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2024)
Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization
von: Qiu, Xinyu, et al.
Veröffentlicht: (2026)
von: Qiu, Xinyu, et al.
Veröffentlicht: (2026)
Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
von: Jin, Jiayun, et al.
Veröffentlicht: (2026)
von: Jin, Jiayun, et al.
Veröffentlicht: (2026)
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
von: Hua, Hang, et al.
Veröffentlicht: (2024)
von: Hua, Hang, et al.
Veröffentlicht: (2024)
Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?
von: Che, Chengan, et al.
Veröffentlicht: (2026)
von: Che, Chengan, et al.
Veröffentlicht: (2026)
SuperCLIP: CLIP with Simple Classification Supervision
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
DriveWorld: 4D Pre-trained Scene Understanding via World Models for Autonomous Driving
von: Min, Chen, et al.
Veröffentlicht: (2024)
von: Min, Chen, et al.
Veröffentlicht: (2024)
DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation
von: Xiong, Zhitong, et al.
Veröffentlicht: (2025)
von: Xiong, Zhitong, et al.
Veröffentlicht: (2025)
A Simple Aerial Detection Baseline of Multimodal Language Models
von: Li, Qingyun, et al.
Veröffentlicht: (2025)
von: Li, Qingyun, et al.
Veröffentlicht: (2025)
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024)
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024)
Enhancing Vision-Language Pre-training with Rich Supervisions
von: Gao, Yuan, et al.
Veröffentlicht: (2024)
von: Gao, Yuan, et al.
Veröffentlicht: (2024)
Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion
von: Miao, Honglei, et al.
Veröffentlicht: (2024)
von: Miao, Honglei, et al.
Veröffentlicht: (2024)
Combating Label Noise With A General Surrogate Model For Sample Selection
von: Liang, Chao, et al.
Veröffentlicht: (2023)
von: Liang, Chao, et al.
Veröffentlicht: (2023)
SVIPTR: Fast and Efficient Scene Text Recognition with Vision Permutable Extractor
von: Cheng, Xianfu, et al.
Veröffentlicht: (2024)
von: Cheng, Xianfu, et al.
Veröffentlicht: (2024)
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
von: Ning, Shan, et al.
Veröffentlicht: (2026)
von: Ning, Shan, et al.
Veröffentlicht: (2026)
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
von: Zohra, Fatimah, et al.
Veröffentlicht: (2025)
von: Zohra, Fatimah, et al.
Veröffentlicht: (2025)
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
von: Fhima, Jonathan, et al.
Veröffentlicht: (2024)
von: Fhima, Jonathan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023) -
MDiff4STR: Mask Diffusion Model for Scene Text Recognition
von: Du, Yongkun, et al.
Veröffentlicht: (2025) -
AudioScenic: Audio-Driven Video Scene Editing
von: Shen, Kaixin, et al.
Veröffentlicht: (2024) -
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2025) -
Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
von: Feng, Tuo, et al.
Veröffentlicht: (2024)