GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
Fuente:
arXiv
Saved in:
| Main Authors: | He, Haibin, Zhang, Jing, Ye, Maoyuan, Liu, Juhua, Du, Bo, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching
by: He, Haibin, et al.
Published: (2024)
by: He, Haibin, et al.
Published: (2024)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)
by: He, Haibin, et al.
Published: (2026)
DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting
by: Ye, Maoyuan, et al.
Published: (2023)
by: Ye, Maoyuan, et al.
Published: (2023)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
by: Ye, Maoyuan, et al.
Published: (2025)
by: Ye, Maoyuan, et al.
Published: (2025)
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
by: Zhang, Xike, et al.
Published: (2026)
by: Zhang, Xike, et al.
Published: (2026)
Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation
by: Ye, Maoyuan, et al.
Published: (2024)
by: Ye, Maoyuan, et al.
Published: (2024)
Adapting Segment Anything Model for Power Transmission Corridor Hazard Segmentation
by: Chen, Hang, et al.
Published: (2025)
by: Chen, Hang, et al.
Published: (2025)
Rethink Sparse Signals for Pose-guided Text-to-image Generation
by: Xuan, Wenjie, et al.
Published: (2025)
by: Xuan, Wenjie, et al.
Published: (2025)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
by: Zhong, Qihuang, et al.
Published: (2026)
by: Zhong, Qihuang, et al.
Published: (2026)
Hear the Scene: Audio-Enhanced Text Spotting
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
by: Xuan, Wenjie, et al.
Published: (2024)
by: Xuan, Wenjie, et al.
Published: (2024)
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
by: Xie, Yu, et al.
Published: (2024)
by: Xie, Yu, et al.
Published: (2024)
On Geometry-Enhanced Parameter-Efficient Fine-Tuning for 3D Scene Segmentation
by: Tang, Liyao, et al.
Published: (2025)
by: Tang, Liyao, et al.
Published: (2025)
Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
PolarMAE: Efficient Fetal Ultrasound Pre-training via Semantic Screening and Polar-Guided Masking
by: Lv, Meng, et al.
Published: (2026)
by: Lv, Meng, et al.
Published: (2026)
DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
by: Shen, Leqi, et al.
Published: (2025)
by: Shen, Leqi, et al.
Published: (2025)
LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text Spotting
by: Su, Yuchen, et al.
Published: (2025)
by: Su, Yuchen, et al.
Published: (2025)
TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition
by: Ye, Xingsong, et al.
Published: (2024)
by: Ye, Xingsong, et al.
Published: (2024)
Match-Stereo-Videos: Bidirectional Alignment for Consistent Dynamic Stereo Matching
by: Jing, Junpeng, et al.
Published: (2024)
by: Jing, Junpeng, et al.
Published: (2024)
Focus Entirety and Perceive Environment for Arbitrary-Shaped Text Detection
by: Han, Xu, et al.
Published: (2024)
by: Han, Xu, et al.
Published: (2024)
MUJICA: Reforming SISR Models for PBR Material Super-Resolution via Cross-Map Attention
by: Du, Xin, et al.
Published: (2025)
by: Du, Xin, et al.
Published: (2025)
Efficiently Leveraging Linguistic Priors for Scene Text Spotting
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
Match Stereo Videos via Bidirectional Alignment
by: Jing, Junpeng, et al.
Published: (2024)
by: Jing, Junpeng, et al.
Published: (2024)
Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
BPDO:Boundary Points Dynamic Optimization for Arbitrary Shape Scene Text Detection
by: Zheng, Jinzhi, et al.
Published: (2024)
by: Zheng, Jinzhi, et al.
Published: (2024)
Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
by: Liu, Xianlin, et al.
Published: (2025)
by: Liu, Xianlin, et al.
Published: (2025)
PoseBench: Benchmarking the Robustness of Pose Estimation Models under Corruptions
by: Ma, Sihan, et al.
Published: (2024)
by: Ma, Sihan, et al.
Published: (2024)
Stereo Any Video: Temporally Consistent Stereo Matching
by: Jing, Junpeng, et al.
Published: (2025)
by: Jing, Junpeng, et al.
Published: (2025)
LOGO: Video Text Spotting with Language Collaboration and Glyph Perception Model
by: Liu, Hongen, et al.
Published: (2024)
by: Liu, Hongen, et al.
Published: (2024)
RFL-CDNet: Towards Accurate Change Detection via Richer Feature Learning
by: Gan, Yuhang, et al.
Published: (2024)
by: Gan, Yuhang, et al.
Published: (2024)
TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
by: Shen, Leqi, et al.
Published: (2024)
by: Shen, Leqi, et al.
Published: (2024)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
by: Xue, Zhucun, et al.
Published: (2025)
by: Xue, Zhucun, et al.
Published: (2025)
DcMatch: Unsupervised Multi-Shape Matching with Dual-Level Consistency
by: Ye, Tianwei, et al.
Published: (2025)
by: Ye, Tianwei, et al.
Published: (2025)
Weak Supervision with Arbitrary Single Frame for Micro- and Macro-expression Spotting
by: Yu, Wang-Wang, et al.
Published: (2024)
by: Yu, Wang-Wang, et al.
Published: (2024)
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
by: Tao, Huaqi, et al.
Published: (2025)
by: Tao, Huaqi, et al.
Published: (2025)
GloTSFormer: Global Video Text Spotting Transformer
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
On Robust Cross-View Consistency in Self-Supervised Monocular Depth Estimation
by: Zhao, Haimei, et al.
Published: (2022)
by: Zhao, Haimei, et al.
Published: (2022)
GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-Mesh
by: Wen, Jing, et al.
Published: (2024)
by: Wen, Jing, et al.
Published: (2024)
Similar Items
-
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching
by: He, Haibin, et al.
Published: (2024) -
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026) -
DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting
by: Ye, Maoyuan, et al.
Published: (2023) -
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
by: He, Haibin, et al.
Published: (2025) -
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
by: Ye, Maoyuan, et al.
Published: (2025)