Saved in:
| Main Authors: | Li, Jing, Wang, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.19504 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficiently Leveraging Linguistic Priors for Scene Text Spotting
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
InstructOCR: Instruction Boosting Scene Text Spotting
by: Duan, Chen, et al.
Published: (2024)
by: Duan, Chen, et al.
Published: (2024)
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
by: Das, Alloy, et al.
Published: (2024)
by: Das, Alloy, et al.
Published: (2024)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
by: Tao, Huaqi, et al.
Published: (2025)
by: Tao, Huaqi, et al.
Published: (2025)
ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting
by: Duan, Chen, et al.
Published: (2024)
by: Duan, Chen, et al.
Published: (2024)
Inverse-like Antagonistic Scene Text Spotting via Reading-Order Estimation and Dynamic Sampling
by: Zhang, Shi-Xue, et al.
Published: (2024)
by: Zhang, Shi-Xue, et al.
Published: (2024)
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
by: Nguyen, Hieu, et al.
Published: (2024)
by: Nguyen, Hieu, et al.
Published: (2024)
GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting
by: Ye, Maoyuan, et al.
Published: (2023)
by: Ye, Maoyuan, et al.
Published: (2023)
Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting
by: Colombo, Antonio, et al.
Published: (2026)
by: Colombo, Antonio, et al.
Published: (2026)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching
by: He, Haibin, et al.
Published: (2024)
by: He, Haibin, et al.
Published: (2024)
Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
by: Liu, Xianlin, et al.
Published: (2025)
by: Liu, Xianlin, et al.
Published: (2025)
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
by: Xie, Yu, et al.
Published: (2024)
by: Xie, Yu, et al.
Published: (2024)
Block-level Text Spotting with LLMs
by: Bannur, Ganesh, et al.
Published: (2024)
by: Bannur, Ganesh, et al.
Published: (2024)
SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance
by: Shen, Guibao, et al.
Published: (2024)
by: Shen, Guibao, et al.
Published: (2024)
Parrot Captions Teach CLIP to Spot Text
by: Lin, Yiqi, et al.
Published: (2023)
by: Lin, Yiqi, et al.
Published: (2023)
GloTSFormer: Global Video Text Spotting Transformer
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
AudioScenic: Audio-Driven Video Scene Editing
by: Shen, Kaixin, et al.
Published: (2024)
by: Shen, Kaixin, et al.
Published: (2024)
Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakes
by: Liu, Weifeng, et al.
Published: (2024)
by: Liu, Weifeng, et al.
Published: (2024)
Text2NeRF: Text-Driven 3D Scene Generation with Neural Radiance Fields
by: Zhang, Jingbo, et al.
Published: (2023)
by: Zhang, Jingbo, et al.
Published: (2023)
AV-Unified: A Unified Framework for Audio-visual Scene Understanding
by: Li, Guangyao, et al.
Published: (2026)
by: Li, Guangyao, et al.
Published: (2026)
Mamba-Enhanced Text-Audio-Video Alignment Network for Emotion Recognition in Conversations
by: Li, Xinran, et al.
Published: (2024)
by: Li, Xinran, et al.
Published: (2024)
Watermark Text Pattern Spotting in Document Images
by: Krubiński, Mateusz, et al.
Published: (2024)
by: Krubiński, Mateusz, et al.
Published: (2024)
Unifying Global and Local Scene Entities Modelling for Precise Action Spotting
by: Tran, Kim Hoang, et al.
Published: (2024)
by: Tran, Kim Hoang, et al.
Published: (2024)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Partial Scene Text Retrieval
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
LOGO: Video Text Spotting with Language Collaboration and Glyph Perception Model
by: Liu, Hongen, et al.
Published: (2024)
by: Liu, Hongen, et al.
Published: (2024)
Hearing Hands: Generating Sounds from Physical Interactions in 3D Scenes
by: Dou, Yiming, et al.
Published: (2025)
by: Dou, Yiming, et al.
Published: (2025)
TextSculptor: Training and Benchmarking Scene Text Editing
by: Lin, Yiheng, et al.
Published: (2026)
by: Lin, Yiheng, et al.
Published: (2026)
Text-Pass Filter: An Efficient Scene Text Detector
by: Yang, Chuang, et al.
Published: (2026)
by: Yang, Chuang, et al.
Published: (2026)
SceneTracker: Long-term Scene Flow Estimation Network
by: Wang, Bo, et al.
Published: (2024)
by: Wang, Bo, et al.
Published: (2024)
MELDAE: A Framework for Micro-Expression Spotting, Detection, and Automatic Evaluation in In-the-Wild Conversational Scenes
by: Feng, Yigui, et al.
Published: (2025)
by: Feng, Yigui, et al.
Published: (2025)
Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
by: Wu, Jianzong, et al.
Published: (2025)
by: Wu, Jianzong, et al.
Published: (2025)
AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting
by: Liu, Yuyuan, et al.
Published: (2025)
by: Liu, Yuyuan, et al.
Published: (2025)
Hyper-Local Deformable Transformers for Text Spotting on Historical Maps
by: Lin, Yijun, et al.
Published: (2025)
by: Lin, Yijun, et al.
Published: (2025)
Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
by: Lee, Junwon, et al.
Published: (2025)
by: Lee, Junwon, et al.
Published: (2025)
Similar Items
-
Efficiently Leveraging Linguistic Priors for Scene Text Spotting
by: Nguyen, Nguyen, et al.
Published: (2024) -
InstructOCR: Instruction Boosting Scene Text Spotting
by: Duan, Chen, et al.
Published: (2024) -
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
by: Huang, Mingxin, et al.
Published: (2024) -
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
by: Das, Alloy, et al.
Published: (2023) -
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
by: Das, Alloy, et al.
Published: (2024)