Advancing Generative Model Evaluation: A Novel Algorithm for Realistic Image Synthesis and Comparison in OCR System
Fuente:
arXiv
Salvato in:
| Autori principali: | Memari, Majid, Ahmed, Khaled R., Rahimi, Shahram, Golilarz, Noorbakhsh Amiri |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Bridging the Gap: Toward Cognitive Autonomy in Artificial Intelligence
di: Golilarz, Noorbakhsh Amiri, et al.
Pubblicazione: (2025)
di: Golilarz, Noorbakhsh Amiri, et al.
Pubblicazione: (2025)
Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
di: Kansana, Manish, et al.
Pubblicazione: (2025)
di: Kansana, Manish, et al.
Pubblicazione: (2025)
Gamma2Patterns: Deep Cognitive Attention Region Identification and Gamma-Alpha Pattern Analysis
di: Jahan, Sobhana, et al.
Pubblicazione: (2026)
di: Jahan, Sobhana, et al.
Pubblicazione: (2026)
One Patch is All You Need: Joint Surface Material Reconstruction and Classification from Minimal Visual Cues
di: Penchala, Sindhuja, et al.
Pubblicazione: (2025)
di: Penchala, Sindhuja, et al.
Pubblicazione: (2025)
Edge-Based Learning for Improved Classification Under Adversarial Noise
di: Kansana, Manish, et al.
Pubblicazione: (2025)
di: Kansana, Manish, et al.
Pubblicazione: (2025)
Where to Bind Matters: Hebbian Fast Weights in Vision Transformers for Few-Shot Character Recognition
di: Money, Gavin, et al.
Pubblicazione: (2026)
di: Money, Gavin, et al.
Pubblicazione: (2026)
Towards Neurocognitive-Inspired Intelligence: From AI's Structural Mimicry to Human-Like Functional Cognition
di: Golilarz, Noorbakhsh Amiri, et al.
Pubblicazione: (2025)
di: Golilarz, Noorbakhsh Amiri, et al.
Pubblicazione: (2025)
Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision
di: Kansana, Manish, et al.
Pubblicazione: (2025)
di: Kansana, Manish, et al.
Pubblicazione: (2025)
Comparison of Image Preprocessing Techniques for Vehicle License Plate Recognition Using OCR: Performance and Accuracy Evaluation
di: Tavares, Renato Augusto
Pubblicazione: (2024)
di: Tavares, Renato Augusto
Pubblicazione: (2024)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
di: Zhang, Junyuan, et al.
Pubblicazione: (2024)
di: Zhang, Junyuan, et al.
Pubblicazione: (2024)
AI Learning Algorithms: Deep Learning, Hybrid Models, and Large-Scale Model Integration
di: Golilarz, Noorbakhsh Amiri, et al.
Pubblicazione: (2024)
di: Golilarz, Noorbakhsh Amiri, et al.
Pubblicazione: (2024)
R-GAT: Cancer Document Classification Leveraging Graph-Based Residual Network for Scenarios with Limited Data
di: Hossain, Elias, et al.
Pubblicazione: (2024)
di: Hossain, Elias, et al.
Pubblicazione: (2024)
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
di: Yuan, Yu, et al.
Pubblicazione: (2024)
di: Yuan, Yu, et al.
Pubblicazione: (2024)
MedInsight: A Multi-Source Context Augmentation Framework for Generating Patient-Centric Medical Responses using Large Language Models
di: Neupane, Subash, et al.
Pubblicazione: (2024)
di: Neupane, Subash, et al.
Pubblicazione: (2024)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
di: Yang, Zhibo, et al.
Pubblicazione: (2024)
di: Yang, Zhibo, et al.
Pubblicazione: (2024)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
di: Chen, Song, et al.
Pubblicazione: (2025)
di: Chen, Song, et al.
Pubblicazione: (2025)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
di: Wei, Haoran, et al.
Pubblicazione: (2024)
di: Wei, Haoran, et al.
Pubblicazione: (2024)
Closer to Ground Truth: Realistic Shape and Appearance Labeled Data Generation for Unsupervised Underwater Image Segmentation
di: Jelea, Andrei, et al.
Pubblicazione: (2025)
di: Jelea, Andrei, et al.
Pubblicazione: (2025)
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
di: Zhang, Peirong, et al.
Pubblicazione: (2025)
di: Zhang, Peirong, et al.
Pubblicazione: (2025)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
di: Shi, Yang, et al.
Pubblicazione: (2025)
di: Shi, Yang, et al.
Pubblicazione: (2025)
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
di: Sun, Lin, et al.
Pubblicazione: (2026)
di: Sun, Lin, et al.
Pubblicazione: (2026)
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
di: Wang, Qilin, et al.
Pubblicazione: (2024)
di: Wang, Qilin, et al.
Pubblicazione: (2024)
Refracting Reality: Generating Images with Realistic Transparent Objects
di: Yin, Yue, et al.
Pubblicazione: (2025)
di: Yin, Yue, et al.
Pubblicazione: (2025)
SCPainter: A Unified Framework for Realistic 3D Asset Insertion and Novel View Synthesis
di: Dobre, Paul, et al.
Pubblicazione: (2025)
di: Dobre, Paul, et al.
Pubblicazione: (2025)
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
di: Zhou, Dewei, et al.
Pubblicazione: (2024)
di: Zhou, Dewei, et al.
Pubblicazione: (2024)
MIRAGE: Model-agnostic Industrial Realistic Anomaly Generation and Evaluation for Visual Anomaly Detection
di: Hu, Jinwei, et al.
Pubblicazione: (2026)
di: Hu, Jinwei, et al.
Pubblicazione: (2026)
Continuous Exposure-Time Modeling for Realistic Atmospheric Turbulence Synthesis
di: Zeng, Junwei, et al.
Pubblicazione: (2026)
di: Zeng, Junwei, et al.
Pubblicazione: (2026)
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
di: Gagnier, Henry, et al.
Pubblicazione: (2026)
di: Gagnier, Henry, et al.
Pubblicazione: (2026)
Text-guided Controllable Diffusion for Realistic Camouflage Images Generation
di: Qian, Yuhang, et al.
Pubblicazione: (2025)
di: Qian, Yuhang, et al.
Pubblicazione: (2025)
Exploring OCR-augmented Generation for Bilingual VQA
di: Lee, JoonHo, et al.
Pubblicazione: (2025)
di: Lee, JoonHo, et al.
Pubblicazione: (2025)
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
di: Xie, Yu, et al.
Pubblicazione: (2025)
di: Xie, Yu, et al.
Pubblicazione: (2025)
OCR-Agent: Agentic OCR with Capability and Memory Reflection
di: Wen, Shimin, et al.
Pubblicazione: (2026)
di: Wen, Shimin, et al.
Pubblicazione: (2026)
OmniOCR: Generalist OCR for Ethnic Minority Languages
di: Liu, Bonan, et al.
Pubblicazione: (2026)
di: Liu, Bonan, et al.
Pubblicazione: (2026)
DODO: Discrete OCR Diffusion Models
di: Man, Sean, et al.
Pubblicazione: (2026)
di: Man, Sean, et al.
Pubblicazione: (2026)
FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
di: Wang, Mengchao, et al.
Pubblicazione: (2025)
di: Wang, Mengchao, et al.
Pubblicazione: (2025)
Realistic Evaluation of Model Merging for Compositional Generalization
di: Tam, Derek, et al.
Pubblicazione: (2024)
di: Tam, Derek, et al.
Pubblicazione: (2024)
A 2-Stage Model for Vehicle Class and Orientation Detection with Photo-Realistic Image Generation
di: Kim, Youngmin, et al.
Pubblicazione: (2025)
di: Kim, Youngmin, et al.
Pubblicazione: (2025)
Prompting Medical Vision-Language Models to Mitigate Diagnosis Bias by Generating Realistic Dermoscopic Images
di: Munia, Nusrat, et al.
Pubblicazione: (2025)
di: Munia, Nusrat, et al.
Pubblicazione: (2025)
From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models
di: Gong, Weile, et al.
Pubblicazione: (2026)
di: Gong, Weile, et al.
Pubblicazione: (2026)
PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation
di: Wu, Fan, et al.
Pubblicazione: (2025)
di: Wu, Fan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Bridging the Gap: Toward Cognitive Autonomy in Artificial Intelligence
di: Golilarz, Noorbakhsh Amiri, et al.
Pubblicazione: (2025) -
Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
di: Kansana, Manish, et al.
Pubblicazione: (2025) -
Gamma2Patterns: Deep Cognitive Attention Region Identification and Gamma-Alpha Pattern Analysis
di: Jahan, Sobhana, et al.
Pubblicazione: (2026) -
One Patch is All You Need: Joint Surface Material Reconstruction and Classification from Minimal Visual Cues
di: Penchala, Sindhuja, et al.
Pubblicazione: (2025) -
Edge-Based Learning for Improved Classification Under Adversarial Noise
di: Kansana, Manish, et al.
Pubblicazione: (2025)