TiCLS : Tightly Coupled Language Text Spotter

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jang, Leeje, Lin, Yijun, Chiang, Yao-Yi, Weinman, Jerod
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908811397169152
author Jang, Leeje
Lin, Yijun
Chiang, Yao-Yi
Weinman, Jerod
author_facet Jang, Leeje
Lin, Yijun
Chiang, Yao-Yi
Weinman, Jerod
contents Scene text spotting aims to detect and recognize text in real-world images, where instances are often short, fragmented, or visually ambiguous. Existing methods primarily rely on visual cues and implicitly capture local character dependencies, but they overlook the benefits of external linguistic knowledge. Prior attempts to integrate language models either adapt language modeling objectives without external knowledge or apply pretrained models that are misaligned with the word-level granularity of scene text. We propose TiCLS, an end-to-end text spotter that explicitly incorporates external linguistic knowledge from a character-level pretrained language model. TiCLS introduces a linguistic decoder that fuses visual and linguistic features, yet can be initialized by a pretrained language model, enabling robust recognition of ambiguous or fragmented text. Experiments on ICDAR 2015 and Total-Text demonstrate that TiCLS achieves state-of-the-art performance, validating the effectiveness of PLM-guided linguistic integration for scene text spotting.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04030
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TiCLS : Tightly Coupled Language Text Spotter
Jang, Leeje
Lin, Yijun
Chiang, Yao-Yi
Weinman, Jerod
Computer Vision and Pattern Recognition
Scene text spotting aims to detect and recognize text in real-world images, where instances are often short, fragmented, or visually ambiguous. Existing methods primarily rely on visual cues and implicitly capture local character dependencies, but they overlook the benefits of external linguistic knowledge. Prior attempts to integrate language models either adapt language modeling objectives without external knowledge or apply pretrained models that are misaligned with the word-level granularity of scene text. We propose TiCLS, an end-to-end text spotter that explicitly incorporates external linguistic knowledge from a character-level pretrained language model. TiCLS introduces a linguistic decoder that fuses visual and linguistic features, yet can be initialized by a pretrained language model, enabling robust recognition of ambiguous or fragmented text. Experiments on ICDAR 2015 and Total-Text demonstrate that TiCLS achieves state-of-the-art performance, validating the effectiveness of PLM-guided linguistic integration for scene text spotting.
title TiCLS : Tightly Coupled Language Text Spotter
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.04030