Saved in:
| Main Authors: | Memari, Majid, Ahmed, Khaled R., Rahimi, Shahram, Golilarz, Noorbakhsh Amiri |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.17204 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Gap: Toward Cognitive Autonomy in Artificial Intelligence
by: Golilarz, Noorbakhsh Amiri, et al.
Published: (2025)
by: Golilarz, Noorbakhsh Amiri, et al.
Published: (2025)
Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
by: Kansana, Manish, et al.
Published: (2025)
by: Kansana, Manish, et al.
Published: (2025)
Gamma2Patterns: Deep Cognitive Attention Region Identification and Gamma-Alpha Pattern Analysis
by: Jahan, Sobhana, et al.
Published: (2026)
by: Jahan, Sobhana, et al.
Published: (2026)
One Patch is All You Need: Joint Surface Material Reconstruction and Classification from Minimal Visual Cues
by: Penchala, Sindhuja, et al.
Published: (2025)
by: Penchala, Sindhuja, et al.
Published: (2025)
Edge-Based Learning for Improved Classification Under Adversarial Noise
by: Kansana, Manish, et al.
Published: (2025)
by: Kansana, Manish, et al.
Published: (2025)
Where to Bind Matters: Hebbian Fast Weights in Vision Transformers for Few-Shot Character Recognition
by: Money, Gavin, et al.
Published: (2026)
by: Money, Gavin, et al.
Published: (2026)
Towards Neurocognitive-Inspired Intelligence: From AI's Structural Mimicry to Human-Like Functional Cognition
by: Golilarz, Noorbakhsh Amiri, et al.
Published: (2025)
by: Golilarz, Noorbakhsh Amiri, et al.
Published: (2025)
Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision
by: Kansana, Manish, et al.
Published: (2025)
by: Kansana, Manish, et al.
Published: (2025)
R-GAT: Cancer Document Classification Leveraging Graph-Based Residual Network for Scenarios with Limited Data
by: Hossain, Elias, et al.
Published: (2024)
by: Hossain, Elias, et al.
Published: (2024)
AI Learning Algorithms: Deep Learning, Hybrid Models, and Large-Scale Model Integration
by: Golilarz, Noorbakhsh Amiri, et al.
Published: (2024)
by: Golilarz, Noorbakhsh Amiri, et al.
Published: (2024)
MedInsight: A Multi-Source Context Augmentation Framework for Generating Patient-Centric Medical Responses using Large Language Models
by: Neupane, Subash, et al.
Published: (2024)
by: Neupane, Subash, et al.
Published: (2024)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
by: Zhang, Junyuan, et al.
Published: (2024)
by: Zhang, Junyuan, et al.
Published: (2024)
Comparison of Image Preprocessing Techniques for Vehicle License Plate Recognition Using OCR: Performance and Accuracy Evaluation
by: Tavares, Renato Augusto
Published: (2024)
by: Tavares, Renato Augusto
Published: (2024)
Estimating Reliability of Electric Vehicle Charging Ecosystem using the Principle of Maximum Entropy
by: Tripathi, Himanshu, et al.
Published: (2025)
by: Tripathi, Himanshu, et al.
Published: (2025)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
by: Yang, Zhibo, et al.
Published: (2024)
by: Yang, Zhibo, et al.
Published: (2024)
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
by: Yuan, Yu, et al.
Published: (2024)
by: Yuan, Yu, et al.
Published: (2024)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
by: Chen, Song, et al.
Published: (2025)
by: Chen, Song, et al.
Published: (2025)
Closer to Ground Truth: Realistic Shape and Appearance Labeled Data Generation for Unsupervised Underwater Image Segmentation
by: Jelea, Andrei, et al.
Published: (2025)
by: Jelea, Andrei, et al.
Published: (2025)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
by: Sun, Lin, et al.
Published: (2026)
by: Sun, Lin, et al.
Published: (2026)
Transfer Learning Applied to Computer Vision Problems: Survey on Current Progress, Limitations, and Opportunities
by: Panda, Aaryan, et al.
Published: (2024)
by: Panda, Aaryan, et al.
Published: (2024)
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation
by: Wang, Qilin, et al.
Published: (2024)
by: Wang, Qilin, et al.
Published: (2024)
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
by: Gagnier, Henry, et al.
Published: (2026)
by: Gagnier, Henry, et al.
Published: (2026)
Refracting Reality: Generating Images with Realistic Transparent Objects
by: Yin, Yue, et al.
Published: (2025)
by: Yin, Yue, et al.
Published: (2025)
SCPainter: A Unified Framework for Realistic 3D Asset Insertion and Novel View Synthesis
by: Dobre, Paul, et al.
Published: (2025)
by: Dobre, Paul, et al.
Published: (2025)
Realistic Evaluation of Model Merging for Compositional Generalization
by: Tam, Derek, et al.
Published: (2024)
by: Tam, Derek, et al.
Published: (2024)
Towards Secure MLOps: Surveying Attacks, Mitigation Strategies, and Research Challenges
by: Patel, Raj, et al.
Published: (2025)
by: Patel, Raj, et al.
Published: (2025)
Continuous Exposure-Time Modeling for Realistic Atmospheric Turbulence Synthesis
by: Zeng, Junwei, et al.
Published: (2026)
by: Zeng, Junwei, et al.
Published: (2026)
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
MIRAGE: Model-agnostic Industrial Realistic Anomaly Generation and Evaluation for Visual Anomaly Detection
by: Hu, Jinwei, et al.
Published: (2026)
by: Hu, Jinwei, et al.
Published: (2026)
Towards 3D Semantic Image Synthesis for Medical Imaging
by: Tang, Wenwu, et al.
Published: (2025)
by: Tang, Wenwu, et al.
Published: (2025)
OCR-Agent: Agentic OCR with Capability and Memory Reflection
by: Wen, Shimin, et al.
Published: (2026)
by: Wen, Shimin, et al.
Published: (2026)
OmniOCR: Generalist OCR for Ethnic Minority Languages
by: Liu, Bonan, et al.
Published: (2026)
by: Liu, Bonan, et al.
Published: (2026)
Exploring OCR-augmented Generation for Bilingual VQA
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
Text-guided Controllable Diffusion for Realistic Camouflage Images Generation
by: Qian, Yuhang, et al.
Published: (2025)
by: Qian, Yuhang, et al.
Published: (2025)
Structuring a Training Strategy to Robustify Perception Models with Realistic Image Augmentations
by: Hammam, Ahmed, et al.
Published: (2024)
by: Hammam, Ahmed, et al.
Published: (2024)
Realistic Noise Synthesis with Diffusion Models
by: Wu, Qi, et al.
Published: (2023)
by: Wu, Qi, et al.
Published: (2023)
DODO: Discrete OCR Diffusion Models
by: Man, Sean, et al.
Published: (2026)
by: Man, Sean, et al.
Published: (2026)
Similar Items
-
Bridging the Gap: Toward Cognitive Autonomy in Artificial Intelligence
by: Golilarz, Noorbakhsh Amiri, et al.
Published: (2025) -
Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
by: Kansana, Manish, et al.
Published: (2025) -
Gamma2Patterns: Deep Cognitive Attention Region Identification and Gamma-Alpha Pattern Analysis
by: Jahan, Sobhana, et al.
Published: (2026) -
One Patch is All You Need: Joint Surface Material Reconstruction and Classification from Minimal Visual Cues
by: Penchala, Sindhuja, et al.
Published: (2025) -
Edge-Based Learning for Improved Classification Under Adversarial Noise
by: Kansana, Manish, et al.
Published: (2025)