Training Kindai OCR with parallel textline images and self-attention feature distance-based loss
Fuente:
arXiv
Saved in:
| Main Authors: | Le, Anh, Kitamoto, Asanobu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Machine Learning for the Digital Typhoon Dataset: Extensions to Multiple Basins and New Developments in Representations and Tasks
by: Kitamoto, Asanobu, et al.
Published: (2024)
by: Kitamoto, Asanobu, et al.
Published: (2024)
Single image super-resolution based on trainable feature matching attention network
by: Chen, Qizhou, et al.
Published: (2024)
by: Chen, Qizhou, et al.
Published: (2024)
Leveraging feature communication in federated learning for remote sensing image classification
by: Duong, Anh-Kiet, et al.
Published: (2024)
by: Duong, Anh-Kiet, et al.
Published: (2024)
Intelligent recognition of GPR road hidden defect images based on feature fusion and attention mechanism
by: Lv, Haotian, et al.
Published: (2025)
by: Lv, Haotian, et al.
Published: (2025)
FALFormer: Feature-aware Landmarks self-attention for Whole-slide Image Classification
by: Bui, Doanh C., et al.
Published: (2024)
by: Bui, Doanh C., et al.
Published: (2024)
Co-distilled attention guided masked image modeling with noisy teacher for self-supervised learning on medical images
by: Jiang, Jue, et al.
Published: (2026)
by: Jiang, Jue, et al.
Published: (2026)
Understanding implementation pitfalls of distance-based metrics for image segmentation
by: Podobnik, Gasper, et al.
Published: (2024)
by: Podobnik, Gasper, et al.
Published: (2024)
OCR-Agent: Agentic OCR with Capability and Memory Reflection
by: Wen, Shimin, et al.
Published: (2026)
by: Wen, Shimin, et al.
Published: (2026)
OmniOCR: Generalist OCR for Ethnic Minority Languages
by: Liu, Bonan, et al.
Published: (2026)
by: Liu, Bonan, et al.
Published: (2026)
SSTFB: Leveraging self-supervised pretext learning and temporal self-attention with feature branching for real-time video polyp segmentation
by: Xu, Ziang, et al.
Published: (2024)
by: Xu, Ziang, et al.
Published: (2024)
Improving underwater semantic segmentation with underwater image quality attention and muti-scale aggregation attention
by: Zuo, Xin, et al.
Published: (2025)
by: Zuo, Xin, et al.
Published: (2025)
3D scene generation from scene graphs and self-attention
by: Bonazzi, Pietro, et al.
Published: (2024)
by: Bonazzi, Pietro, et al.
Published: (2024)
YOLO algorithm with hybrid attention feature pyramid network for solder joint defect detection
by: Ang, Li, et al.
Published: (2024)
by: Ang, Li, et al.
Published: (2024)
Zero-shot sketch-based remote sensing image retrieval based on multi-level and attention-guided tokenization
by: Yang, Bo, et al.
Published: (2024)
by: Yang, Bo, et al.
Published: (2024)
Agentar-Fin-OCR
by: Qian, Siyi, et al.
Published: (2026)
by: Qian, Siyi, et al.
Published: (2026)
An explainable hierarchical self attention-based approach for tremor detection in the time domain
by: Odonga, Timothy, et al.
Published: (2026)
by: Odonga, Timothy, et al.
Published: (2026)
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
by: Sun, Lin, et al.
Published: (2026)
by: Sun, Lin, et al.
Published: (2026)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
by: Zhang, Junyuan, et al.
Published: (2024)
by: Zhang, Junyuan, et al.
Published: (2024)
Real-time estimation of overt attention from dynamic features of the face using deep-learning
by: Ortubay, Aimar Silvan, et al.
Published: (2024)
by: Ortubay, Aimar Silvan, et al.
Published: (2024)
Unpaired Image Dehazing via Kolmogorov-Arnold Transformation of Latent Features
by: Tran, Le-Anh
Published: (2025)
by: Tran, Le-Anh
Published: (2025)
Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
by: Tien, Dong Nguyen, et al.
Published: (2025)
by: Tien, Dong Nguyen, et al.
Published: (2025)
Trinity Detector:text-assisted and attention mechanisms based spectral fusion for diffusion generation image detection
by: Song, Jiawei, et al.
Published: (2024)
by: Song, Jiawei, et al.
Published: (2024)
Towards a text-based quantitative and explainable histopathology image analysis
by: Nguyen, Anh Tien, et al.
Published: (2024)
by: Nguyen, Anh Tien, et al.
Published: (2024)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
by: Chen, Song, et al.
Published: (2025)
by: Chen, Song, et al.
Published: (2025)
TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content
by: Anand, Avinash, et al.
Published: (2024)
by: Anand, Avinash, et al.
Published: (2024)
olmOCR 2: Unit Test Rewards for Document OCR
by: Poznanski, Jake, et al.
Published: (2025)
by: Poznanski, Jake, et al.
Published: (2025)
Exploring PCA-based feature representations of image pixels via CNN to enhance food image segmentation
by: Dai, Ying
Published: (2024)
by: Dai, Ying
Published: (2024)
Segmentation-free Connectionist Temporal Classification loss based OCR Model for Text Captcha Classification
by: Khatavkar, Vaibhav, et al.
Published: (2024)
by: Khatavkar, Vaibhav, et al.
Published: (2024)
A self-attention model for robust rigid slice-to-volume registration of functional MRI
by: Khawaled, Samah, et al.
Published: (2024)
by: Khawaled, Samah, et al.
Published: (2024)
Efficient feature matching for UAV images based on compact GPU data scheduling
by: Jiang, San, et al.
Published: (2025)
by: Jiang, San, et al.
Published: (2025)
Polaffini: A feature-based approach for robust affine and polyaffine image registration
by: Legouhy, Antoine, et al.
Published: (2026)
by: Legouhy, Antoine, et al.
Published: (2026)
ABot-OCR Technical Report
by: Jiang, Kaitao, et al.
Published: (2026)
by: Jiang, Kaitao, et al.
Published: (2026)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
by: Yang, Zhibo, et al.
Published: (2024)
by: Yang, Zhibo, et al.
Published: (2024)
Probabilistic smooth attention for deep multiple instance learning in medical imaging
by: Castro-Macías, Francisco M., et al.
Published: (2025)
by: Castro-Macías, Francisco M., et al.
Published: (2025)
Joint multi-dimensional dynamic attention and transformer for general image restoration
by: Zhang, Huan, et al.
Published: (2024)
by: Zhang, Huan, et al.
Published: (2024)
Multi-Contrast Fusion Module: An attention mechanism integrating multi-contrast features for fetal torso plane classification
by: Zhu, Shengjun, et al.
Published: (2025)
by: Zhu, Shengjun, et al.
Published: (2025)
Local positional graphs and attentive local features for a data and runtime-efficient hierarchical place recognition pipeline
by: Yuan, Fangming, et al.
Published: (2024)
by: Yuan, Fangming, et al.
Published: (2024)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
Error Patterns in Historical OCR: A Comparative Analysis of TrOCR and a Vision-Language Model
by: Vesalainen, Ari, et al.
Published: (2026)
by: Vesalainen, Ari, et al.
Published: (2026)
Similar Items
-
Machine Learning for the Digital Typhoon Dataset: Extensions to Multiple Basins and New Developments in Representations and Tasks
by: Kitamoto, Asanobu, et al.
Published: (2024) -
Single image super-resolution based on trainable feature matching attention network
by: Chen, Qizhou, et al.
Published: (2024) -
Leveraging feature communication in federated learning for remote sensing image classification
by: Duong, Anh-Kiet, et al.
Published: (2024) -
Intelligent recognition of GPR road hidden defect images based on feature fusion and attention mechanism
by: Lv, Haotian, et al.
Published: (2025) -
FALFormer: Feature-aware Landmarks self-attention for Whole-slide Image Classification
by: Bui, Doanh C., et al.
Published: (2024)