What's Wrong with the Bottom-up Methods in Arbitrary-shape Scene Text Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Chengpei, Jia, Wenjing, Cui, Tingcheng, Wang, Ruomei, Zhang, Yuan-fang, He, Xiangjian |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MorphText: Deep Morphology Regularized Arbitrary-shape Scene Text Detection
by: Xu, Chengpei, et al.
Published: (2024)
by: Xu, Chengpei, et al.
Published: (2024)
Seeing Text in the Dark: Algorithm and Benchmark
by: Xu, Chengpei, et al.
Published: (2024)
by: Xu, Chengpei, et al.
Published: (2024)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
by: Zhao, Baoquan, et al.
Published: (2025)
by: Zhao, Baoquan, et al.
Published: (2025)
PixelatedScatter: Arbitrary-level Visual Abstraction for Large-scale Multiclass Scatterplots
by: Guo, Ziheng, et al.
Published: (2025)
by: Guo, Ziheng, et al.
Published: (2025)
Scene-Text Grounding for Text-Based Video Question Answering
by: Zhou, Sheng, et al.
Published: (2024)
by: Zhou, Sheng, et al.
Published: (2024)
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
by: Zhao, Yuan, et al.
Published: (2026)
by: Zhao, Yuan, et al.
Published: (2026)
Bridging Discrete and Continuous: A Multimodal Strategy for Complex Emotion Detection
by: Jia, Jiehui, et al.
Published: (2024)
by: Jia, Jiehui, et al.
Published: (2024)
PFB-Diff: Progressive Feature Blending Diffusion for Text-driven Image Editing
by: Huang, Wenjing, et al.
Published: (2023)
by: Huang, Wenjing, et al.
Published: (2023)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
by: Zhao, Lingzhi, et al.
Published: (2025)
by: Zhao, Lingzhi, et al.
Published: (2025)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
by: Meng, Yiran, et al.
Published: (2025)
by: Meng, Yiran, et al.
Published: (2025)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal Fusion
by: Gong, Ziyu, et al.
Published: (2024)
by: Gong, Ziyu, et al.
Published: (2024)
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
by: Dong, Wenqi, et al.
Published: (2025)
by: Dong, Wenqi, et al.
Published: (2025)
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
by: Wang, Yongqi, et al.
Published: (2025)
by: Wang, Yongqi, et al.
Published: (2025)
HumanVLM: Foundation for Human-Scene Vision-Language Model
by: Dai, Dawei, et al.
Published: (2024)
by: Dai, Dawei, et al.
Published: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval
by: Qin, Yawen, et al.
Published: (2026)
by: Qin, Yawen, et al.
Published: (2026)
Prototypical Prompting for Text-to-image Person Re-identification
by: Yan, Shuanglin, et al.
Published: (2024)
by: Yan, Shuanglin, et al.
Published: (2024)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
by: Tan, Jiawei, et al.
Published: (2024)
by: Tan, Jiawei, et al.
Published: (2024)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
by: Zhang, Hanlei, et al.
Published: (2024)
by: Zhang, Hanlei, et al.
Published: (2024)
STEFANN: Scene Text Editor using Font Adaptive Neural Network
by: Roy, Prasun, et al.
Published: (2019)
by: Roy, Prasun, et al.
Published: (2019)
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration
by: Xie, Tianyidan, et al.
Published: (2026)
by: Xie, Tianyidan, et al.
Published: (2026)
Deep joint source-channel coding for wireless point cloud transmission
by: Zhang, Cixiao, et al.
Published: (2024)
by: Zhang, Cixiao, et al.
Published: (2024)
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
by: Wang, Yuyue, et al.
Published: (2025)
by: Wang, Yuyue, et al.
Published: (2025)
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
by: Xu, Junhao, et al.
Published: (2025)
by: Xu, Junhao, et al.
Published: (2025)
OpenVNA: A Framework for Analyzing the Behavior of Multimodal Language Understanding System under Noisy Scenarios
by: Yuan, Ziqi, et al.
Published: (2024)
by: Yuan, Ziqi, et al.
Published: (2024)
RFNNS: Robust Fixed Neural Network Steganography with Universal Text-to-Image Models
by: Cheng, Yu, et al.
Published: (2025)
by: Cheng, Yu, et al.
Published: (2025)
Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning
by: Wang, Hanyao, et al.
Published: (2024)
by: Wang, Hanyao, et al.
Published: (2024)
TeleAntiFraud-28k: An Audio-Text Slow-Thinking Dataset for Telecom Fraud Detection
by: Ma, Zhiming, et al.
Published: (2025)
by: Ma, Zhiming, et al.
Published: (2025)
Efficient Video to Audio Mapper with Visual Scene Detection
by: Yi, Mingjing, et al.
Published: (2024)
by: Yi, Mingjing, et al.
Published: (2024)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
by: Li, Qingcao, et al.
Published: (2026)
by: Li, Qingcao, et al.
Published: (2026)
TF-Mamba: Text-enhanced Fusion Mamba with Missing Modalities for Robust Multimodal Sentiment Analysis
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
by: Yuan, Shenghai, et al.
Published: (2024)
by: Yuan, Shenghai, et al.
Published: (2024)
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait Recognition
by: Li, Xinzhu, et al.
Published: (2025)
by: Li, Xinzhu, et al.
Published: (2025)
Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description
by: Pan, Mianzhi, et al.
Published: (2023)
by: Pan, Mianzhi, et al.
Published: (2023)
Similar Items
-
MorphText: Deep Morphology Regularized Arbitrary-shape Scene Text Detection
by: Xu, Chengpei, et al.
Published: (2024) -
Seeing Text in the Dark: Algorithm and Benchmark
by: Xu, Chengpei, et al.
Published: (2024) -
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
by: Zhao, Baoquan, et al.
Published: (2025) -
PixelatedScatter: Arbitrary-level Visual Abstraction for Large-scale Multiclass Scatterplots
by: Guo, Ziheng, et al.
Published: (2025) -
Scene-Text Grounding for Text-Based Video Question Answering
by: Zhou, Sheng, et al.
Published: (2024)