Saved in:
| Main Authors: | Bai, Weikang, Du, Yongkun, Su, Yuchen, Xie, Yazhen, Chen, Zhineng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.13731 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
Instruction-Guided Scene Text Recognition
by: Du, Yongkun, et al.
Published: (2024)
by: Du, Yongkun, et al.
Published: (2024)
TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition
by: Ye, Xingsong, et al.
Published: (2024)
by: Ye, Xingsong, et al.
Published: (2024)
LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text Spotting
by: Su, Yuchen, et al.
Published: (2025)
by: Su, Yuchen, et al.
Published: (2025)
Decoder Pre-Training with only Text for Scene Text Recognition
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition
by: Du, Yongkun, et al.
Published: (2024)
by: Du, Yongkun, et al.
Published: (2024)
Explicit Relational Reasoning Network for Scene Text Detection
by: Su, Yuchen, et al.
Published: (2024)
by: Su, Yuchen, et al.
Published: (2024)
What Is Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution
by: Ye, Xingsong, et al.
Published: (2026)
by: Ye, Xingsong, et al.
Published: (2026)
Bidirectional Trained Tree-Structured Decoder for Handwritten Mathematical Expression Recognition
by: Cheng, Hanbo, et al.
Published: (2023)
by: Cheng, Hanbo, et al.
Published: (2023)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Towards Scalable Training for Handwritten Mathematical Expression Recognition
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Out of Length Text Recognition with Sub-String Matching
by: Du, Yongkun, et al.
Published: (2024)
by: Du, Yongkun, et al.
Published: (2024)
ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
by: Felemban, Abdulwahab, et al.
Published: (2025)
by: Felemban, Abdulwahab, et al.
Published: (2025)
SemiHMER: Semi-supervised Handwritten Mathematical Expression Recognition using pseudo-labels
by: Chen, Kehua, et al.
Published: (2025)
by: Chen, Kehua, et al.
Published: (2025)
PointCloud-Text Matching: Benchmark Datasets and a Baseline
by: Feng, Yanglin, et al.
Published: (2024)
by: Feng, Yanglin, et al.
Published: (2024)
Evaluating Facial Expression Recognition Datasets for Deep Learning: A Benchmark Study with Novel Similarity Metrics
by: Gaya-Morey, F. Xavier, et al.
Published: (2025)
by: Gaya-Morey, F. Xavier, et al.
Published: (2025)
MathNet: A Data-Centric Approach for Printed Mathematical Expression Recognition
by: Schmitt-Koopmann, Felix M., et al.
Published: (2024)
by: Schmitt-Koopmann, Felix M., et al.
Published: (2024)
Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
by: Xing, Rui, et al.
Published: (2025)
by: Xing, Rui, et al.
Published: (2025)
MDiff4STR: Mask Diffusion Model for Scene Text Recognition
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
by: Liang, Baoyu, et al.
Published: (2025)
by: Liang, Baoyu, et al.
Published: (2025)
Segment Anything for Satellite Imagery: A Strong Baseline and a Regional Dataset for Automatic Field Delineation
by: Scribano, Carmelo, et al.
Published: (2025)
by: Scribano, Carmelo, et al.
Published: (2025)
Auditing Facial Emotion Recognition Datasets for Posed Expressions and Racial Bias
by: Khan, Rina, et al.
Published: (2025)
by: Khan, Rina, et al.
Published: (2025)
Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
by: Li, Haiyang, et al.
Published: (2025)
by: Li, Haiyang, et al.
Published: (2025)
A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
by: Jiang, Siyang, et al.
Published: (2025)
by: Jiang, Siyang, et al.
Published: (2025)
UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark
by: Zhou, Zhaokun, et al.
Published: (2024)
by: Zhou, Zhaokun, et al.
Published: (2024)
MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
by: Zhao, Kaikai, et al.
Published: (2025)
by: Zhao, Kaikai, et al.
Published: (2025)
Compound Expression Recognition via Large Vision-Language Models
by: Yu, Jun, et al.
Published: (2025)
by: Yu, Jun, et al.
Published: (2025)
3DCoMPaT$^{++}$: An improved Large-scale 3D Vision Dataset for Compositional Recognition
by: Slim, Habib, et al.
Published: (2023)
by: Slim, Habib, et al.
Published: (2023)
TAG: Thinking with Action Unit Grounding for Facial Expression Recognition
by: Lin, Haobo, et al.
Published: (2026)
by: Lin, Haobo, et al.
Published: (2026)
StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
by: Simonyan, Aleksandr, et al.
Published: (2026)
by: Simonyan, Aleksandr, et al.
Published: (2026)
SimpleMatch: A Simple and Strong Baseline for Semantic Correspondence
by: Jin, Hailing, et al.
Published: (2026)
by: Jin, Hailing, et al.
Published: (2026)
Adaptive Global-Local Representation Learning and Selection for Cross-Domain Facial Expression Recognition
by: Gao, Yuefang, et al.
Published: (2024)
by: Gao, Yuefang, et al.
Published: (2024)
Strong Baseline: Multi-UAV Tracking via YOLOv12 with BoT-SORT-ReID
by: Chen, Yu-Hsi
Published: (2025)
by: Chen, Yu-Hsi
Published: (2025)
InjectFlow: Weak Guides Strong via Orthogonal Injection for Flow Matching
by: Wang, Dayu, et al.
Published: (2026)
by: Wang, Dayu, et al.
Published: (2026)
RAW: Robust Avatar Watermarking -- Benchmarking and Baseline
by: Parry, Jack, et al.
Published: (2026)
by: Parry, Jack, et al.
Published: (2026)
Signal-SGN++: Topology-Enhanced Time-Frequency Spiking Graph Network for Skeleton-Based Action Recognition
by: Zheng, Naichuan, et al.
Published: (2025)
by: Zheng, Naichuan, et al.
Published: (2025)
High Performance Space Debris Tracking in Complex Skylight Backgrounds with a Large-Scale Dataset
by: Zhuang, Guohang, et al.
Published: (2025)
by: Zhuang, Guohang, et al.
Published: (2025)
Generalizable Facial Expression Recognition
by: Zhang, Yuhang, et al.
Published: (2024)
by: Zhang, Yuhang, et al.
Published: (2024)
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
by: Yao, Xincheng, et al.
Published: (2026)
by: Yao, Xincheng, et al.
Published: (2026)
Similar Items
-
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
by: Du, Yongkun, et al.
Published: (2025) -
UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
by: Du, Yongkun, et al.
Published: (2025) -
Instruction-Guided Scene Text Recognition
by: Du, Yongkun, et al.
Published: (2024) -
TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition
by: Ye, Xingsong, et al.
Published: (2024) -
LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text Spotting
by: Su, Yuchen, et al.
Published: (2025)