Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Bin, Wu, Fan, Ouyang, Linke, Gu, Zhuangcheng, Zhang, Rui, Xia, Renqiu, Zhang, Bo, He, Conghui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
LOCR: Location-Guided Transformer for Optical Character Recognition
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
by: Kang, Hengrui, et al.
Published: (2025)
by: Kang, Hengrui, et al.
Published: (2025)
Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
by: Zhao, Zhiyuan, et al.
Published: (2023)
by: Zhao, Zhiyuan, et al.
Published: (2023)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
by: Zhang, Junyuan, et al.
Published: (2024)
by: Zhang, Junyuan, et al.
Published: (2024)
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
by: Yan, Xiangchao, et al.
Published: (2025)
by: Yan, Xiangchao, et al.
Published: (2025)
DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
by: Zhang, Qintong, et al.
Published: (2025)
by: Zhang, Qintong, et al.
Published: (2025)
Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition
by: Zhou, Bangbang, et al.
Published: (2024)
by: Zhou, Bangbang, et al.
Published: (2024)
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
by: Xiong, Minhao, et al.
Published: (2025)
by: Xiong, Minhao, et al.
Published: (2025)
Zero-Shot Chinese Character Recognition with Hierarchical Multi-Granularity Image-Text Aligning
by: Zhu, Yinglian, et al.
Published: (2025)
by: Zhu, Yinglian, et al.
Published: (2025)
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
Multi-Modal Character Localization and Extraction for Chinese Text Recognition
by: Li, Qilong, et al.
Published: (2026)
by: Li, Qilong, et al.
Published: (2026)
Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
by: Ye, Hancheng, et al.
Published: (2024)
by: Ye, Hancheng, et al.
Published: (2024)
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
by: Dong, Xiaoyi, et al.
Published: (2024)
by: Dong, Xiaoyi, et al.
Published: (2024)
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
by: Liao, Ning, et al.
Published: (2023)
by: Liao, Ning, et al.
Published: (2023)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization
by: Luo, Yucong, et al.
Published: (2024)
by: Luo, Yucong, et al.
Published: (2024)
MatchDet: A Collaborative Framework for Image Matching and Object Detection
by: Lai, Jinxiang, et al.
Published: (2023)
by: Lai, Jinxiang, et al.
Published: (2023)
Revealing and Mitigating the Challenge of Detecting Character Knowledge Errors in LLM Role-Playing
by: Zhang, Wenyuan, et al.
Published: (2024)
by: Zhang, Wenyuan, et al.
Published: (2024)
GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
by: Feng, Yuan, et al.
Published: (2025)
by: Feng, Yuan, et al.
Published: (2025)
Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition
by: Lin, Tiancheng, et al.
Published: (2024)
by: Lin, Tiancheng, et al.
Published: (2024)
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
by: Tu, Quan, et al.
Published: (2024)
by: Tu, Quan, et al.
Published: (2024)
DEMO: A Statistical Perspective for Efficient Image-Text Matching
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Image-to-LaTeX Converter for Mathematical Formulas and Text
by: Gurgurov, Daniil, et al.
Published: (2024)
by: Gurgurov, Daniil, et al.
Published: (2024)
Text-Region Matching for Multi-Label Image Recognition with Missing Labels
by: Ma, Leilei, et al.
Published: (2024)
by: Ma, Leilei, et al.
Published: (2024)
OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
by: Huang, Qidong, et al.
Published: (2023)
by: Huang, Qidong, et al.
Published: (2023)
MinerU: An Open-Source Solution for Precise Document Content Extraction
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Image Copy-Move Forgery Detection via Deep PatchMatch and Pairwise Ranking Learning
by: Li, Yuanman, et al.
Published: (2024)
by: Li, Yuanman, et al.
Published: (2024)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
Comateformer: Combined Attention Transformer for Semantic Sentence Matching
by: Li, Bo, et al.
Published: (2024)
by: Li, Bo, et al.
Published: (2024)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
by: Yue, Xinli, et al.
Published: (2025)
by: Yue, Xinli, et al.
Published: (2025)
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
by: Cui, Wanqing, et al.
Published: (2024)
by: Cui, Wanqing, et al.
Published: (2024)
AMR-CCR: Anchored Modular Retrieval for Continual Chinese Character Recognition
by: Wu, Yuchuan, et al.
Published: (2026)
by: Wu, Yuchuan, et al.
Published: (2026)
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching
by: He, Haibin, et al.
Published: (2024)
by: He, Haibin, et al.
Published: (2024)
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
by: Beňová, Ivana, et al.
Published: (2024)
by: Beňová, Ivana, et al.
Published: (2024)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
by: Zhang, Huixuan, et al.
Published: (2025)
by: Zhang, Huixuan, et al.
Published: (2025)
ComCLIP: Training-Free Compositional Image and Text Matching
by: Jiang, Kenan, et al.
Published: (2022)
by: Jiang, Kenan, et al.
Published: (2022)
Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
by: Hu, Li, et al.
Published: (2023)
by: Hu, Li, et al.
Published: (2023)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
Similar Items
-
UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition
by: Wang, Bin, et al.
Published: (2024) -
LOCR: Location-Guided Transformer for Optical Character Recognition
by: Sun, Yu, et al.
Published: (2024) -
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
by: Kang, Hengrui, et al.
Published: (2025) -
Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
by: Zhao, Zhiyuan, et al.
Published: (2023) -
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
by: Zhang, Junyuan, et al.
Published: (2024)