DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Maoyuan, Zhang, Jing, Zhao, Shanshan, Liu, Juhua, Liu, Tongliang, Du, Bo, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching
by: He, Haibin, et al.
Published: (2024)
by: He, Haibin, et al.
Published: (2024)
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
by: Zhang, Xike, et al.
Published: (2026)
by: Zhang, Xike, et al.
Published: (2026)
Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation
by: Ye, Maoyuan, et al.
Published: (2024)
by: Ye, Maoyuan, et al.
Published: (2024)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)
by: He, Haibin, et al.
Published: (2026)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
Rethink Sparse Signals for Pose-guided Text-to-image Generation
by: Xuan, Wenjie, et al.
Published: (2025)
by: Xuan, Wenjie, et al.
Published: (2025)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
by: Ye, Maoyuan, et al.
Published: (2025)
by: Ye, Maoyuan, et al.
Published: (2025)
When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
by: Xuan, Wenjie, et al.
Published: (2024)
by: Xuan, Wenjie, et al.
Published: (2024)
Adapting Segment Anything Model for Power Transmission Corridor Hazard Segmentation
by: Chen, Hang, et al.
Published: (2025)
by: Chen, Hang, et al.
Published: (2025)
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
by: Zhong, Qihuang, et al.
Published: (2026)
by: Zhong, Qihuang, et al.
Published: (2026)
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning
by: Chen, Yiyang, et al.
Published: (2025)
by: Chen, Yiyang, et al.
Published: (2025)
Local-consistent Transformation Learning for Rotation-invariant Point Cloud Analysis
by: Chen, Yiyang, et al.
Published: (2024)
by: Chen, Yiyang, et al.
Published: (2024)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
by: Das, Alloy, et al.
Published: (2024)
by: Das, Alloy, et al.
Published: (2024)
Hear the Scene: Audio-Enhanced Text Spotting
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
ProtoSolo: Interpretable Image Classification via Single-Prototype Activation
by: Peng, Yitao, et al.
Published: (2025)
by: Peng, Yitao, et al.
Published: (2025)
UniMix: Towards Domain Adaptive and Generalizable LiDAR Semantic Segmentation in Adverse Weather
by: Zhao, Haimei, et al.
Published: (2024)
by: Zhao, Haimei, et al.
Published: (2024)
SynAgent: Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy
by: Yao, Wei, et al.
Published: (2026)
by: Yao, Wei, et al.
Published: (2026)
SimDistill: Simulated Multi-modal Distillation for BEV 3D Object Detection
by: Zhao, Haimei, et al.
Published: (2023)
by: Zhao, Haimei, et al.
Published: (2023)
Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
by: Liu, Xianlin, et al.
Published: (2025)
by: Liu, Xianlin, et al.
Published: (2025)
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
by: Lu, Runnan, et al.
Published: (2025)
by: Lu, Runnan, et al.
Published: (2025)
SoloParkour: Constrained Reinforcement Learning for Visual Locomotion from Privileged Experience
by: Chane-Sane, Elliot, et al.
Published: (2024)
by: Chane-Sane, Elliot, et al.
Published: (2024)
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
PolarMAE: Efficient Fetal Ultrasound Pre-training via Semantic Screening and Polar-Guided Masking
by: Lv, Meng, et al.
Published: (2026)
by: Lv, Meng, et al.
Published: (2026)
RFL-CDNet: Towards Accurate Change Detection via Richer Feature Learning
by: Gan, Yuhang, et al.
Published: (2024)
by: Gan, Yuhang, et al.
Published: (2024)
Dual-disentangled Deep Multiple Clustering
by: Yao, Jiawei, et al.
Published: (2024)
by: Yao, Jiawei, et al.
Published: (2024)
CogMorph: Cognitive Morphing Attacks for Text-to-Image Models
by: Jing, Zonglei, et al.
Published: (2025)
by: Jing, Zonglei, et al.
Published: (2025)
When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety Guidance
by: Xiang, Yongli, et al.
Published: (2026)
by: Xiang, Yongli, et al.
Published: (2026)
On Robust Cross-View Consistency in Self-Supervised Monocular Depth Estimation
by: Zhao, Haimei, et al.
Published: (2022)
by: Zhao, Haimei, et al.
Published: (2022)
Towards Modality-agnostic Label-efficient Segmentation with Entropy-Regularized Distribution Alignment
by: Tang, Liyao, et al.
Published: (2024)
by: Tang, Liyao, et al.
Published: (2024)
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
Architecture, Dataset and Model-Scale Agnostic Data-free Meta-Learning
by: Hu, Zixuan, et al.
Published: (2023)
by: Hu, Zixuan, et al.
Published: (2023)
Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
by: Cao, Min, et al.
Published: (2025)
by: Cao, Min, et al.
Published: (2025)
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)
by: Lei, Shiye, et al.
Published: (2023)
Deep Learning-Based Point Cloud Registration: A Comprehensive Survey and Taxonomy
by: Zhang, Yu-Xin, et al.
Published: (2024)
by: Zhang, Yu-Xin, et al.
Published: (2024)
MUJICA: Reforming SISR Models for PBR Material Super-Resolution via Cross-Map Attention
by: Du, Xin, et al.
Published: (2025)
by: Du, Xin, et al.
Published: (2025)
Decoder Pre-Training with only Text for Scene Text Recognition
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
SubFlow: Sub-mode Conditioned Flow Matching for Diverse One-Step Generation
by: Lin, Yexiong, et al.
Published: (2026)
by: Lin, Yexiong, et al.
Published: (2026)
AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing
by: Hong, Ziming, et al.
Published: (2025)
by: Hong, Ziming, et al.
Published: (2025)
Similar Items
-
GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
by: He, Haibin, et al.
Published: (2025) -
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching
by: He, Haibin, et al.
Published: (2024) -
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
by: Zhang, Xike, et al.
Published: (2026) -
Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation
by: Ye, Maoyuan, et al.
Published: (2024) -
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)