Efficiently Leveraging Linguistic Priors for Scene Text Spotting
Fuente:
arXiv
Guardado en:
| Autores principales: | Nguyen, Nguyen, Tian, Yapeng, Xu, Chenliang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
por: Nguyen, Nhi Ngoc-Yen, et al.
Publicado: (2026)
por: Nguyen, Nhi Ngoc-Yen, et al.
Publicado: (2026)
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
por: Nguyen, Hieu, et al.
Publicado: (2024)
por: Nguyen, Hieu, et al.
Publicado: (2024)
OSCaR: Object State Captioning and State Change Representation
por: Nguyen, Nguyen, et al.
Publicado: (2024)
por: Nguyen, Nguyen, et al.
Publicado: (2024)
Scaling Concept With Text-Guided Diffusion Models
por: Huang, Chao, et al.
Publicado: (2024)
por: Huang, Chao, et al.
Publicado: (2024)
Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
por: Liang, Susan, et al.
Publicado: (2024)
por: Liang, Susan, et al.
Publicado: (2024)
Hear the Scene: Audio-Enhanced Text Spotting
por: Li, Jing, et al.
Publicado: (2024)
por: Li, Jing, et al.
Publicado: (2024)
High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
por: Huang, Chao, et al.
Publicado: (2025)
por: Huang, Chao, et al.
Publicado: (2025)
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
por: Das, Alloy, et al.
Publicado: (2023)
por: Das, Alloy, et al.
Publicado: (2023)
Leveraging Geometric Priors for Unaligned Scene Change Detection
por: Liu, Ziling, et al.
Publicado: (2025)
por: Liu, Ziling, et al.
Publicado: (2025)
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
por: Huang, Mingxin, et al.
Publicado: (2024)
por: Huang, Mingxin, et al.
Publicado: (2024)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
por: Das, Alloy, et al.
Publicado: (2024)
por: Das, Alloy, et al.
Publicado: (2024)
FreSca: Scaling in Frequency Space Enhances Diffusion Models
por: Huang, Chao, et al.
Publicado: (2025)
por: Huang, Chao, et al.
Publicado: (2025)
Learning Dynamic Scene Reconstruction with Sinusoidal Geometric Priors
por: Guo, Tian, et al.
Publicado: (2025)
por: Guo, Tian, et al.
Publicado: (2025)
Smart Camera Parking System With Auto Parking Spot Detection
por: Nguyen, Tuan T., et al.
Publicado: (2024)
por: Nguyen, Tuan T., et al.
Publicado: (2024)
Inverse-like Antagonistic Scene Text Spotting via Reading-Order Estimation and Dynamic Sampling
por: Zhang, Shi-Xue, et al.
Publicado: (2024)
por: Zhang, Shi-Xue, et al.
Publicado: (2024)
Text-Pass Filter: An Efficient Scene Text Detector
por: Yang, Chuang, et al.
Publicado: (2026)
por: Yang, Chuang, et al.
Publicado: (2026)
InstructOCR: Instruction Boosting Scene Text Spotting
por: Duan, Chen, et al.
Publicado: (2024)
por: Duan, Chen, et al.
Publicado: (2024)
High-Quality Visually-Guided Sound Separation from Diverse Categories
por: Huang, Chao, et al.
Publicado: (2023)
por: Huang, Chao, et al.
Publicado: (2023)
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
por: Nguyen, Nghia Hieu, et al.
Publicado: (2024)
por: Nguyen, Nghia Hieu, et al.
Publicado: (2024)
ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting
por: Duan, Chen, et al.
Publicado: (2024)
por: Duan, Chen, et al.
Publicado: (2024)
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
por: Zeng, Weichao, et al.
Publicado: (2024)
por: Zeng, Weichao, et al.
Publicado: (2024)
Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting
por: Colombo, Antonio, et al.
Publicado: (2026)
por: Colombo, Antonio, et al.
Publicado: (2026)
Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
por: Zhang, Yifei, et al.
Publicado: (2025)
por: Zhang, Yifei, et al.
Publicado: (2025)
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
por: Huang, Chao, et al.
Publicado: (2025)
por: Huang, Chao, et al.
Publicado: (2025)
Leveraging Habitat Information for Fine-grained Bird Identification
por: Nguyen, Tin, et al.
Publicado: (2023)
por: Nguyen, Tin, et al.
Publicado: (2023)
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
por: Tao, Huaqi, et al.
Publicado: (2025)
por: Tao, Huaqi, et al.
Publicado: (2025)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
por: Li, Wenrui, et al.
Publicado: (2024)
por: Li, Wenrui, et al.
Publicado: (2024)
Leveraging Text Localization for Scene Text Removal via Text-aware Masked Image Modeling
por: Wang, Zixiao, et al.
Publicado: (2024)
por: Wang, Zixiao, et al.
Publicado: (2024)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
por: Lyu, Jiahao, et al.
Publicado: (2024)
por: Lyu, Jiahao, et al.
Publicado: (2024)
NCAP: Scene Text Image Super-Resolution with Non-CAtegorical Prior
por: Park, Dongwoo, et al.
Publicado: (2025)
por: Park, Dongwoo, et al.
Publicado: (2025)
HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding
por: Nguyen, Trong-Thuan, et al.
Publicado: (2023)
por: Nguyen, Trong-Thuan, et al.
Publicado: (2023)
EAGLE: Egocentric AGgregated Language-video Engine
por: Bi, Jing, et al.
Publicado: (2024)
por: Bi, Jing, et al.
Publicado: (2024)
From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
por: Li, Jia, et al.
Publicado: (2025)
por: Li, Jia, et al.
Publicado: (2025)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
por: Shu, Yan, et al.
Publicado: (2025)
por: Shu, Yan, et al.
Publicado: (2025)
Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model Guidance
por: Pham, Duc-Hai, et al.
Publicado: (2024)
por: Pham, Duc-Hai, et al.
Publicado: (2024)
LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text Spotting
por: Su, Yuchen, et al.
Publicado: (2025)
por: Su, Yuchen, et al.
Publicado: (2025)
Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
por: Tourani, Siddharth, et al.
Publicado: (2025)
por: Tourani, Siddharth, et al.
Publicado: (2025)
Text-IRSTD: Leveraging Semantic Text to Promote Infrared Small Target Detection in Complex Scenes
por: Huang, Feng, et al.
Publicado: (2025)
por: Huang, Feng, et al.
Publicado: (2025)
Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition
por: Le, Kha Nhat, et al.
Publicado: (2024)
por: Le, Kha Nhat, et al.
Publicado: (2024)
SATURN: Autoregressive Image Generation Guided by Scene Graphs
por: Vo, Thanh-Nhan, et al.
Publicado: (2025)
por: Vo, Thanh-Nhan, et al.
Publicado: (2025)
Ejemplares similares
-
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
por: Nguyen, Nhi Ngoc-Yen, et al.
Publicado: (2026) -
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
por: Nguyen, Hieu, et al.
Publicado: (2024) -
OSCaR: Object State Captioning and State Change Representation
por: Nguyen, Nguyen, et al.
Publicado: (2024) -
Scaling Concept With Text-Guided Diffusion Models
por: Huang, Chao, et al.
Publicado: (2024) -
Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
por: Liang, Susan, et al.
Publicado: (2024)