A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
Fuente:
arXiv
Saved in:
| Main Authors: | Gomez, Lluis, Karatzas, Dimosthenis |
|---|---|
| Format: | Preprint |
| Published: |
2014
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
Reading in the Dark: Low-light Scene Text Recognition
by: Fu, Xuanshuo, et al.
Published: (2026)
by: Fu, Xuanshuo, et al.
Published: (2026)
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
by: Mishra, Pritam, et al.
Published: (2025)
by: Mishra, Pritam, et al.
Published: (2025)
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
Reading Between the Lanes: Text VideoQA on the Road
by: Tom, George, et al.
Published: (2023)
by: Tom, George, et al.
Published: (2023)
Machine Unlearning for Document Classification
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
by: Mishra, Pritam, et al.
Published: (2026)
by: Mishra, Pritam, et al.
Published: (2026)
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
by: Li, Zongmin, et al.
Published: (2026)
by: Li, Zongmin, et al.
Published: (2026)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
by: Pintore, Marco, et al.
Published: (2025)
by: Pintore, Marco, et al.
Published: (2025)
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
by: Ortega, Marc Serra, et al.
Published: (2025)
by: Ortega, Marc Serra, et al.
Published: (2025)
GRIF-DM: Generation of Rich Impression Fonts using Diffusion Models
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
Image-text matching for large-scale book collections
by: Llabrés, Artemis, et al.
Published: (2024)
by: Llabrés, Artemis, et al.
Published: (2024)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
One missing piece in Vision and Language: A Survey on Comics Understanding
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
ComicsPAP: understanding comic strips by picking the correct panel
by: Vivoli, Emanuele, et al.
Published: (2025)
by: Vivoli, Emanuele, et al.
Published: (2025)
MorphText: Deep Morphology Regularized Arbitrary-shape Scene Text Detection
by: Xu, Chengpei, et al.
Published: (2024)
by: Xu, Chengpei, et al.
Published: (2024)
BPDO:Boundary Points Dynamic Optimization for Arbitrary Shape Scene Text Detection
by: Zheng, Jinzhi, et al.
Published: (2024)
by: Zheng, Jinzhi, et al.
Published: (2024)
CSSL-MHTR: Continual Self-Supervised Learning for Scalable Multi-script Handwritten Text Recognition
by: Dhiaf, Marwa, et al.
Published: (2023)
by: Dhiaf, Marwa, et al.
Published: (2023)
FastScene: Text-Driven Fast 3D Indoor Scene Generation via Panoramic Gaussian Splatting
by: Ma, Yikun, et al.
Published: (2024)
by: Ma, Yikun, et al.
Published: (2024)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
by: Das, Alloy, et al.
Published: (2024)
by: Das, Alloy, et al.
Published: (2024)
Text Modality Oriented Image Feature Extraction for Detecting Diffusion-based DeepFake
by: Yang, Di, et al.
Published: (2024)
by: Yang, Di, et al.
Published: (2024)
SVIPTR: Fast and Efficient Scene Text Recognition with Vision Permutable Extractor
by: Cheng, Xianfu, et al.
Published: (2024)
by: Cheng, Xianfu, et al.
Published: (2024)
Language-Driven Object-Oriented Two-Stage Method for Scene Graph Anticipation
by: Zhu, Xiaomeng, et al.
Published: (2025)
by: Zhu, Xiaomeng, et al.
Published: (2025)
Text2Traffic: A Text-to-Image Generation and Editing Method for Traffic Scenes
by: Lv, Feng, et al.
Published: (2025)
by: Lv, Feng, et al.
Published: (2025)
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
Fast SceneScript: Fast and Accurate Language-Based 3D Scene Understanding via Multi-Token Prediction
by: Yin, Ruihong, et al.
Published: (2025)
by: Yin, Ruihong, et al.
Published: (2025)
HSM: Hierarchical Scene Motifs for Multi-Scale Indoor Scene Generation
by: Pun, Hou In Derek, et al.
Published: (2025)
by: Pun, Hou In Derek, et al.
Published: (2025)
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
by: Xie, Yu, et al.
Published: (2024)
by: Xie, Yu, et al.
Published: (2024)
Hierarchical Context Transformer for Multi-level Semantic Scene Understanding
by: Hao, Luoying, et al.
Published: (2025)
by: Hao, Luoying, et al.
Published: (2025)
EUFCC-CIR: a Composed Image Retrieval Dataset for GLAM Collections
by: Net, Francesc, et al.
Published: (2024)
by: Net, Francesc, et al.
Published: (2024)
Towards Arbitrary Motion Completing via Hierarchical Continuous Representation
by: Xu, Chenghao, et al.
Published: (2025)
by: Xu, Chenghao, et al.
Published: (2025)
Task-wise Sampling Convolutions for Arbitrary-Oriented Object Detection in Aerial Images
by: Huang, Zhanchao, et al.
Published: (2022)
by: Huang, Zhanchao, et al.
Published: (2022)
MBQuant: A Novel Multi-Branch Topology Method for Arbitrary Bit-width Network Quantization
by: Zhong, Yunshan, et al.
Published: (2023)
by: Zhong, Yunshan, et al.
Published: (2023)
Similar Items
-
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
by: Kang, Lei, et al.
Published: (2025) -
Reading in the Dark: Low-light Scene Text Recognition
by: Fu, Xuanshuo, et al.
Published: (2026) -
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
by: Vivoli, Emanuele, et al.
Published: (2024) -
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
by: Mishra, Pritam, et al.
Published: (2025) -
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024)