FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
Fuente:
arXiv
Saved in:
| Main Authors: | Das, Alloy, Biswas, Sanket, Pal, Umapada, Lladós, Josep, Bhattacharya, Saumik |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
STEFANN: Scene Text Editor using Font Adaptive Neural Network
by: Roy, Prasun, et al.
Published: (2019)
by: Roy, Prasun, et al.
Published: (2019)
DocRevive: A Unified Pipeline for Document Text Restoration
by: Purkayastha, Kunal, et al.
Published: (2026)
by: Purkayastha, Kunal, et al.
Published: (2026)
TIPS: Text-Induced Pose Synthesis
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
Correlation Weighted Prototype-based Self-Supervised One-Shot Segmentation of Medical Images
by: Manna, Siladittya, et al.
Published: (2024)
by: Manna, Siladittya, et al.
Published: (2024)
CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
MIO : Mutual Information Optimization using Self-Supervised Binary Contrastive Learning
by: Manna, Siladittya, et al.
Published: (2021)
by: Manna, Siladittya, et al.
Published: (2021)
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition
by: Tiwari, Adarsh, et al.
Published: (2024)
by: Tiwari, Adarsh, et al.
Published: (2024)
Scene Aware Person Image Generation through Global Contextual Conditioning
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering
by: Banerjee, Ayan, et al.
Published: (2025)
by: Banerjee, Ayan, et al.
Published: (2025)
DRG-Font: Dynamic Reference-Guided Few-shot Font Generation via Contrastive Style-Content Disentanglement
by: Chakraborty, Rejoy, et al.
Published: (2026)
by: Chakraborty, Rejoy, et al.
Published: (2026)
Multi-scale Attention Guided Pose Transfer
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
GeoContrastNet: Contrastive Key-Value Edge Learning for Language-Agnostic Document Understanding
by: Biescas, Nil, et al.
Published: (2024)
by: Biescas, Nil, et al.
Published: (2024)
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
by: Xie, Yu, et al.
Published: (2024)
by: Xie, Yu, et al.
Published: (2024)
Effects of Degradations on Deep Neural Network Architectures
by: Roy, Prasun, et al.
Published: (2018)
by: Roy, Prasun, et al.
Published: (2018)
Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
by: Chattopadhyay, Soumitri, et al.
Published: (2024)
by: Chattopadhyay, Soumitri, et al.
Published: (2024)
Decorrelation-based Self-Supervised Visual Representation Learning for Writer Identification
by: Maitra, Arkadip, et al.
Published: (2024)
by: Maitra, Arkadip, et al.
Published: (2024)
Position and Rotation Invariant Sign Language Recognition from 3D Kinect Data with Recurrent Neural Networks
by: Roy, Prasun, et al.
Published: (2020)
by: Roy, Prasun, et al.
Published: (2020)
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
by: Roy, Prasun, et al.
Published: (2025)
by: Roy, Prasun, et al.
Published: (2025)
Semantically Consistent Person Image Generation
by: Roy, Prasun, et al.
Published: (2023)
by: Roy, Prasun, et al.
Published: (2023)
EdgeSpotter: Multi-Scale Dense Text Spotting for Industrial Panel Monitoring
by: Fu, Changhong, et al.
Published: (2025)
by: Fu, Changhong, et al.
Published: (2025)
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
Dynamically Scaled Temperature in Self-Supervised Contrastive Learning
by: Manna, Siladittya, et al.
Published: (2023)
by: Manna, Siladittya, et al.
Published: (2023)
Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
TiCLS : Tightly Coupled Language Text Spotter
by: Jang, Leeje, et al.
Published: (2026)
by: Jang, Leeje, et al.
Published: (2026)
Spotter+GPT: Turning Sign Spottings into Sentences with LLMs
by: Sincan, Ozge Mercanoglu, et al.
Published: (2024)
by: Sincan, Ozge Mercanoglu, et al.
Published: (2024)
d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
by: Roy, Prasun, et al.
Published: (2025)
by: Roy, Prasun, et al.
Published: (2025)
TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
by: Zhai, Yukun, et al.
Published: (2023)
by: Zhai, Yukun, et al.
Published: (2023)
GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
by: Bakkali, Souhail, et al.
Published: (2023)
by: Bakkali, Souhail, et al.
Published: (2023)
Hear the Scene: Audio-Enhanced Text Spotting
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting
by: Ye, Maoyuan, et al.
Published: (2023)
by: Ye, Maoyuan, et al.
Published: (2023)
Efficiently Leveraging Linguistic Priors for Scene Text Spotting
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
SAGE-GAN: Towards Realistic and Robust Segmentation of Spatially Ordered Nanoparticles via Attention-Guided GANs
by: Pal, Anindya, et al.
Published: (2026)
by: Pal, Anindya, et al.
Published: (2026)
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
by: Xie, Yu, et al.
Published: (2025)
by: Xie, Yu, et al.
Published: (2025)
Simba: Mamba augmented U-ShiftGCN for Skeletal Action Recognition in Videos
by: Chaudhuri, Soumyabrata, et al.
Published: (2024)
by: Chaudhuri, Soumyabrata, et al.
Published: (2024)
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
Similar Items
-
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
by: Das, Alloy, et al.
Published: (2023) -
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
by: Das, Alloy, et al.
Published: (2023) -
GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
by: Banerjee, Ayan, et al.
Published: (2024) -
STEFANN: Scene Text Editor using Font Adaptive Neural Network
by: Roy, Prasun, et al.
Published: (2019) -
DocRevive: A Unified Pipeline for Document Text Restoration
by: Purkayastha, Kunal, et al.
Published: (2026)