SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Mingxin, Peng, Dezhi, Li, Hongliang, Peng, Zhenghao, Liu, Chongyu, Lin, Dahua, Liu, Yuliang, Bai, Xiang, Jin, Lianwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
von: Peng, Dezhi, et al.
Veröffentlicht: (2023)
von: Peng, Dezhi, et al.
Veröffentlicht: (2023)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
von: Liu, Yuliang, et al.
Veröffentlicht: (2024)
Progressive Evolution from Single-Point to Polygon for Scene Text
von: Deng, Linger, et al.
Veröffentlicht: (2023)
von: Deng, Linger, et al.
Veröffentlicht: (2023)
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
von: Das, Alloy, et al.
Veröffentlicht: (2024)
von: Das, Alloy, et al.
Veröffentlicht: (2024)
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
von: Xie, Yu, et al.
Veröffentlicht: (2024)
von: Xie, Yu, et al.
Veröffentlicht: (2024)
Predicting the Original Appearance of Damaged Historical Documents
von: Yang, Zhenhua, et al.
Veröffentlicht: (2024)
von: Yang, Zhenhua, et al.
Veröffentlicht: (2024)
EdgeSpotter: Multi-Scale Dense Text Spotting for Industrial Panel Monitoring
von: Fu, Changhong, et al.
Veröffentlicht: (2025)
von: Fu, Changhong, et al.
Veröffentlicht: (2025)
UPOCR: Towards Unified Pixel-Level OCR Interface
von: Peng, Dezhi, et al.
Veröffentlicht: (2023)
von: Peng, Dezhi, et al.
Veröffentlicht: (2023)
LEGO: Self-Supervised Representation Learning for Scene Text Images
von: Ren, Yujin, et al.
Veröffentlicht: (2024)
von: Ren, Yujin, et al.
Veröffentlicht: (2024)
Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
TextSleuth: Towards Explainable Tampered Text Detection
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
The First Swahili Language Scene Text Detection and Recognition Dataset
von: Douamba, Fadila Wendigoundi, et al.
Veröffentlicht: (2024)
von: Douamba, Fadila Wendigoundi, et al.
Veröffentlicht: (2024)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
von: Li, Hongliang, et al.
Veröffentlicht: (2025)
von: Li, Hongliang, et al.
Veröffentlicht: (2025)
TiCLS : Tightly Coupled Language Text Spotter
von: Jang, Leeje, et al.
Veröffentlicht: (2026)
von: Jang, Leeje, et al.
Veröffentlicht: (2026)
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
von: Yin, Liang, et al.
Veröffentlicht: (2025)
von: Yin, Liang, et al.
Veröffentlicht: (2025)
Revisiting Tampered Scene Text Detection in the Era of Generative AI
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
von: Qu, Chenfan, et al.
Veröffentlicht: (2024)
Privacy-Preserving Biometric Verification with Handwritten Random Digit String
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
C$^{3}$Bench: A Comprehensive Classical Chinese Understanding Benchmark for Large Language Models
von: Cao, Jiahuan, et al.
Veröffentlicht: (2024)
von: Cao, Jiahuan, et al.
Veröffentlicht: (2024)
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
von: Fu, Ling, et al.
Veröffentlicht: (2024)
von: Fu, Ling, et al.
Veröffentlicht: (2024)
HierCode: A Lightweight Hierarchical Codebook for Zero-shot Chinese Text Recognition
von: Zhang, Yuyi, et al.
Veröffentlicht: (2024)
von: Zhang, Yuyi, et al.
Veröffentlicht: (2024)
SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
von: Luo, Dongliang, et al.
Veröffentlicht: (2025)
von: Luo, Dongliang, et al.
Veröffentlicht: (2025)
OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
von: Wan, Jianqiang, et al.
Veröffentlicht: (2024)
von: Wan, Jianqiang, et al.
Veröffentlicht: (2024)
Datasets for Large Language Models: A Comprehensive Survey
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Spotter+GPT: Turning Sign Spottings into Sentences with LLMs
von: Sincan, Ozge Mercanoglu, et al.
Veröffentlicht: (2024)
von: Sincan, Ozge Mercanoglu, et al.
Veröffentlicht: (2024)
DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
von: Liao, Wenhui, et al.
Veröffentlicht: (2024)
von: Liao, Wenhui, et al.
Veröffentlicht: (2024)
Partial Scene Text Retrieval
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
Toward Real Text Manipulation Detection: New Dataset and New Solution
von: Luo, Dongliang, et al.
Veröffentlicht: (2023)
von: Luo, Dongliang, et al.
Veröffentlicht: (2023)
TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
von: Zhai, Yukun, et al.
Veröffentlicht: (2023)
von: Zhai, Yukun, et al.
Veröffentlicht: (2023)
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
von: Tao, Huaqi, et al.
Veröffentlicht: (2025)
von: Tao, Huaqi, et al.
Veröffentlicht: (2025)
Towards Training-Free Scene Text Editing
von: Li, Yubo, et al.
Veröffentlicht: (2026)
von: Li, Yubo, et al.
Veröffentlicht: (2026)
Smaller But Better: Unifying Layout Generation with Smaller Large Language Models
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
von: Maryam, Hiba, et al.
Veröffentlicht: (2024)
von: Maryam, Hiba, et al.
Veröffentlicht: (2024)
URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding
von: Shi, Yongxin, et al.
Veröffentlicht: (2025)
von: Shi, Yongxin, et al.
Veröffentlicht: (2025)
TextShield-R1: Reinforced Reasoning for Tampered Text Detection
von: Qu, Chenfan, et al.
Veröffentlicht: (2026)
von: Qu, Chenfan, et al.
Veröffentlicht: (2026)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
Hear the Scene: Audio-Enhanced Text Spotting
von: Li, Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
von: Peng, Dezhi, et al.
Veröffentlicht: (2023) -
Bridging the Gap Between End-to-End and Two-Step Text Spotting
von: Huang, Mingxin, et al.
Veröffentlicht: (2024) -
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
von: Liu, Yuliang, et al.
Veröffentlicht: (2024) -
Progressive Evolution from Single-Point to Polygon for Scene Text
von: Deng, Linger, et al.
Veröffentlicht: (2023) -
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)