Entropy Heat-Mapping: Localizing GPT-Based OCR Errors with Sliding-Window Shannon Analysis
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Kaltchenko, Alexei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Assessing GPT Model Uncertainty in Mathematical OCR Tasks via Entropy Analysis
von: Kaltchenko, Alexei
Veröffentlicht: (2024)
von: Kaltchenko, Alexei
Veröffentlicht: (2024)
Slide-SAM: Medical SAM Meets Sliding Window
von: Quan, Quan, et al.
Veröffentlicht: (2023)
von: Quan, Quan, et al.
Veröffentlicht: (2023)
Error Patterns in Historical OCR: A Comparative Analysis of TrOCR and a Vision-Language Model
von: Vesalainen, Ari, et al.
Veröffentlicht: (2026)
von: Vesalainen, Ari, et al.
Veröffentlicht: (2026)
AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising
von: Cui, Liyuan, et al.
Veröffentlicht: (2026)
von: Cui, Liyuan, et al.
Veröffentlicht: (2026)
Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs
von: Ding, Xuan, et al.
Veröffentlicht: (2025)
von: Ding, Xuan, et al.
Veröffentlicht: (2025)
SWinGS: Sliding Windows for Dynamic 3D Gaussian Splatting
von: Shaw, Richard, et al.
Veröffentlicht: (2023)
von: Shaw, Richard, et al.
Veröffentlicht: (2023)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
von: Shi, Yang, et al.
Veröffentlicht: (2025)
von: Shi, Yang, et al.
Veröffentlicht: (2025)
Confidence-Aware Document OCR Error Detection
von: Hemmer, Arthur, et al.
Veröffentlicht: (2024)
von: Hemmer, Arthur, et al.
Veröffentlicht: (2024)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
von: Kashid, Harshvivek, et al.
Veröffentlicht: (2024)
von: Kashid, Harshvivek, et al.
Veröffentlicht: (2024)
SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attribution
von: Wang, Chao, et al.
Veröffentlicht: (2026)
von: Wang, Chao, et al.
Veröffentlicht: (2026)
Sliding Window Attention for Learned Video Compression
von: Kopte, Alexander, et al.
Veröffentlicht: (2025)
von: Kopte, Alexander, et al.
Veröffentlicht: (2025)
Dual Frequency Branch Framework with Reconstructed Sliding Windows Attention for AI-Generated Image Detection
von: Yan, Jiazhen, et al.
Veröffentlicht: (2025)
von: Yan, Jiazhen, et al.
Veröffentlicht: (2025)
WSD-MIL: Window Scale Decay Multiple Instance Learning for Whole Slide Image Classification
von: Feng, Le, et al.
Veröffentlicht: (2025)
von: Feng, Le, et al.
Veröffentlicht: (2025)
SWAT: Sliding Window Adversarial Training for Gradual Domain Adaptation
von: Wang, Zixi, et al.
Veröffentlicht: (2025)
von: Wang, Zixi, et al.
Veröffentlicht: (2025)
OCR-Agent: Agentic OCR with Capability and Memory Reflection
von: Wen, Shimin, et al.
Veröffentlicht: (2026)
von: Wen, Shimin, et al.
Veröffentlicht: (2026)
OmniOCR: Generalist OCR for Ethnic Minority Languages
von: Liu, Bonan, et al.
Veröffentlicht: (2026)
von: Liu, Bonan, et al.
Veröffentlicht: (2026)
SWiT-4D: Sliding-Window Transformer for Lossless and Parameter-Free Temporal 4D Generation
von: Gong, Kehong, et al.
Veröffentlicht: (2025)
von: Gong, Kehong, et al.
Veröffentlicht: (2025)
Continual Multiple Instance Learning with Enhanced Localization for Histopathological Whole Slide Image Analysis
von: Lee, Byung Hyun, et al.
Veröffentlicht: (2025)
von: Lee, Byung Hyun, et al.
Veröffentlicht: (2025)
Reading Between the Lines: Abstaining from VLM-Generated OCR Errors via Latent Representation Probes
von: Yao, Jihan, et al.
Veröffentlicht: (2025)
von: Yao, Jihan, et al.
Veröffentlicht: (2025)
SwinGS: Sliding Window Gaussian Splatting for Volumetric Video Streaming with Arbitrary Length
von: Liu, Bangya, et al.
Veröffentlicht: (2024)
von: Liu, Bangya, et al.
Veröffentlicht: (2024)
The Character Error Vector: Decomposable errors for page-level OCR evaluation
von: Bourne, Jonathan, et al.
Veröffentlicht: (2026)
von: Bourne, Jonathan, et al.
Veröffentlicht: (2026)
Impact of Localization Errors on Label Quality for Online HD Map Construction
von: Blumberg, Alexander, et al.
Veröffentlicht: (2026)
von: Blumberg, Alexander, et al.
Veröffentlicht: (2026)
FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation
von: Wu, Yunfeng, et al.
Veröffentlicht: (2025)
von: Wu, Yunfeng, et al.
Veröffentlicht: (2025)
Agentar-Fin-OCR
von: Qian, Siyi, et al.
Veröffentlicht: (2026)
von: Qian, Siyi, et al.
Veröffentlicht: (2026)
Slide-based Graph Collaborative Training for Histopathology Whole Slide Image Analysis
von: Shi, Jun, et al.
Veröffentlicht: (2024)
von: Shi, Jun, et al.
Veröffentlicht: (2024)
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
von: Sun, Lin, et al.
Veröffentlicht: (2026)
von: Sun, Lin, et al.
Veröffentlicht: (2026)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)
Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory
von: Deng, Tianchen, et al.
Veröffentlicht: (2026)
von: Deng, Tianchen, et al.
Veröffentlicht: (2026)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
von: Chen, Song, et al.
Veröffentlicht: (2025)
von: Chen, Song, et al.
Veröffentlicht: (2025)
TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
olmOCR 2: Unit Test Rewards for Document OCR
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
von: Poznanski, Jake, et al.
Veröffentlicht: (2025)
AEM: Attention Entropy Maximization for Multiple Instance Learning based Whole Slide Image Classification
von: Zhang, Yunlong, et al.
Veröffentlicht: (2024)
von: Zhang, Yunlong, et al.
Veröffentlicht: (2024)
ABot-OCR Technical Report
von: Jiang, Kaitao, et al.
Veröffentlicht: (2026)
von: Jiang, Kaitao, et al.
Veröffentlicht: (2026)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
Underground Mapping and Localization Based on Ground-Penetrating Radar
von: Zhang, Jinchang, et al.
Veröffentlicht: (2024)
von: Zhang, Jinchang, et al.
Veröffentlicht: (2024)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
von: Wei, Haoran, et al.
Veröffentlicht: (2024)
von: Wei, Haoran, et al.
Veröffentlicht: (2024)
DODO: Discrete OCR Diffusion Models
von: Man, Sean, et al.
Veröffentlicht: (2026)
von: Man, Sean, et al.
Veröffentlicht: (2026)
An Empirical Study of Scaling Law for OCR
von: Rang, Miao, et al.
Veröffentlicht: (2023)
von: Rang, Miao, et al.
Veröffentlicht: (2023)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
von: He, Haibin, et al.
Veröffentlicht: (2025)
von: He, Haibin, et al.
Veröffentlicht: (2025)
LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR
von: Taghadouini, Said, et al.
Veröffentlicht: (2026)
von: Taghadouini, Said, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Assessing GPT Model Uncertainty in Mathematical OCR Tasks via Entropy Analysis
von: Kaltchenko, Alexei
Veröffentlicht: (2024) -
Slide-SAM: Medical SAM Meets Sliding Window
von: Quan, Quan, et al.
Veröffentlicht: (2023) -
Error Patterns in Historical OCR: A Comparative Analysis of TrOCR and a Vision-Language Model
von: Vesalainen, Ari, et al.
Veröffentlicht: (2026) -
AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising
von: Cui, Liyuan, et al.
Veröffentlicht: (2026) -
Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs
von: Ding, Xuan, et al.
Veröffentlicht: (2025)