Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments
Fuente:
arXiv
Salvato in:
| Autori principali: | Nagaonkar, Sankalp, Sharma, Augustya, Choithani, Ashish, Trivedi, Ashutosh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
di: Sharma, Shivam, et al.
Pubblicazione: (2026)
di: Sharma, Shivam, et al.
Pubblicazione: (2026)
BadScan: An Architectural Backdoor Attack on Visual State Space Models
di: Deshmukh, Om Suhas, et al.
Pubblicazione: (2024)
di: Deshmukh, Om Suhas, et al.
Pubblicazione: (2024)
TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering
di: Rajkhowa, Tonmoy, et al.
Pubblicazione: (2024)
di: Rajkhowa, Tonmoy, et al.
Pubblicazione: (2024)
PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language
di: Haq, Ijazul, et al.
Pubblicazione: (2025)
di: Haq, Ijazul, et al.
Pubblicazione: (2025)
Invizo: Arabic Handwritten Document Optical Character Recognition Solution
di: Waly, Alhossien, et al.
Pubblicazione: (2025)
di: Waly, Alhossien, et al.
Pubblicazione: (2025)
Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification
di: Addepalli, Sravanti, et al.
Pubblicazione: (2023)
di: Addepalli, Sravanti, et al.
Pubblicazione: (2023)
Focus on the Whole Character: Discriminative Character Modeling for Scene Text Recognition
di: Zhou, Bangbang, et al.
Pubblicazione: (2024)
di: Zhou, Bangbang, et al.
Pubblicazione: (2024)
LaVPR: Benchmarking Language and Vision for Place Recognition
di: Idan, Ofer, et al.
Pubblicazione: (2026)
di: Idan, Ofer, et al.
Pubblicazione: (2026)
Logios : An open source Greek Polytonic Optical Character Recognition system
di: Konstantinos, Perifanos, et al.
Pubblicazione: (2025)
di: Konstantinos, Perifanos, et al.
Pubblicazione: (2025)
GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
Rethinking Genomic Modeling Through Optical Character Recognition
di: Xiang, Hongxin, et al.
Pubblicazione: (2026)
di: Xiang, Hongxin, et al.
Pubblicazione: (2026)
LOCR: Location-Guided Transformer for Optical Character Recognition
di: Sun, Yu, et al.
Pubblicazione: (2024)
di: Sun, Yu, et al.
Pubblicazione: (2024)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
di: Sharma, Saurav, et al.
Pubblicazione: (2025)
di: Sharma, Saurav, et al.
Pubblicazione: (2025)
Nepali Sign Language Characters Recognition: Dataset Development and Deep Learning Approaches
di: Poudel, Birat, et al.
Pubblicazione: (2025)
di: Poudel, Birat, et al.
Pubblicazione: (2025)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
di: Peng, Jingwei, et al.
Pubblicazione: (2025)
di: Peng, Jingwei, et al.
Pubblicazione: (2025)
Words as Geometric Features: Estimating Homography using Optical Character Recognition as Compressed Image Representation
di: Greer, Ross, et al.
Pubblicazione: (2025)
di: Greer, Ross, et al.
Pubblicazione: (2025)
Low-Resource Heuristics for Bahnaric Optical Character Recognition Improvement
di: Tran, Phat, et al.
Pubblicazione: (2026)
di: Tran, Phat, et al.
Pubblicazione: (2026)
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
di: Waseda, Futa, et al.
Pubblicazione: (2025)
di: Waseda, Futa, et al.
Pubblicazione: (2025)
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
di: Ray, Sourjyadip, et al.
Pubblicazione: (2024)
di: Ray, Sourjyadip, et al.
Pubblicazione: (2024)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
di: Darabi, Nastaran, et al.
Pubblicazione: (2026)
di: Darabi, Nastaran, et al.
Pubblicazione: (2026)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
di: Hong, Wenyi, et al.
Pubblicazione: (2025)
di: Hong, Wenyi, et al.
Pubblicazione: (2025)
Breaking the Barriers: Video Vision Transformers for Word-Level Sign Language Recognition
di: Brettmann, Alexander, et al.
Pubblicazione: (2025)
di: Brettmann, Alexander, et al.
Pubblicazione: (2025)
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
di: Chen, Xiuyuan, et al.
Pubblicazione: (2023)
di: Chen, Xiuyuan, et al.
Pubblicazione: (2023)
Character Mixing for Video Generation
di: Liao, Tingting, et al.
Pubblicazione: (2025)
di: Liao, Tingting, et al.
Pubblicazione: (2025)
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
di: Bhatia, Gagan, et al.
Pubblicazione: (2024)
di: Bhatia, Gagan, et al.
Pubblicazione: (2024)
The Urban Vision Hackathon Dataset and Models: Towards Image Annotations and Accurate Vision Models for Indian Traffic
di: Sharma, Akash, et al.
Pubblicazione: (2025)
di: Sharma, Akash, et al.
Pubblicazione: (2025)
Benchmarking Large Language Models for Handwritten Text Recognition
di: Crosilla, Giorgia, et al.
Pubblicazione: (2025)
di: Crosilla, Giorgia, et al.
Pubblicazione: (2025)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
di: Xu, Zhenlin, et al.
Pubblicazione: (2023)
di: Xu, Zhenlin, et al.
Pubblicazione: (2023)
Rotate Your Character: Revisiting Video Diffusion Models for High-Quality 3D Character Generation
di: Wang, Jin, et al.
Pubblicazione: (2026)
di: Wang, Jin, et al.
Pubblicazione: (2026)
Shape and Texture Recognition in Large Vision-Language Models
di: Eppel, Sagi, et al.
Pubblicazione: (2025)
di: Eppel, Sagi, et al.
Pubblicazione: (2025)
Open-Set Recognition in the Age of Vision-Language Models
di: Miller, Dimity, et al.
Pubblicazione: (2024)
di: Miller, Dimity, et al.
Pubblicazione: (2024)
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving
di: Gopalkrishnan, Akshay, et al.
Pubblicazione: (2024)
di: Gopalkrishnan, Akshay, et al.
Pubblicazione: (2024)
SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
di: Du, Hao, et al.
Pubblicazione: (2025)
di: Du, Hao, et al.
Pubblicazione: (2025)
AnimationBench: Are Video Models Good at Character-Centric Animation?
di: Wu, Leyi, et al.
Pubblicazione: (2026)
di: Wu, Leyi, et al.
Pubblicazione: (2026)
MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling
di: Men, Yifang, et al.
Pubblicazione: (2024)
di: Men, Yifang, et al.
Pubblicazione: (2024)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
di: Sathe, Ashutosh, et al.
Pubblicazione: (2024)
di: Sathe, Ashutosh, et al.
Pubblicazione: (2024)
Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
di: Bharadwaj, Siddhant, et al.
Pubblicazione: (2026)
di: Bharadwaj, Siddhant, et al.
Pubblicazione: (2026)
MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis
di: Qiu, Di, et al.
Pubblicazione: (2024)
di: Qiu, Di, et al.
Pubblicazione: (2024)
Distilling Vision-Language Models on Millions of Videos
di: Zhao, Yue, et al.
Pubblicazione: (2024)
di: Zhao, Yue, et al.
Pubblicazione: (2024)
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos
di: Li, Kaining, et al.
Pubblicazione: (2025)
di: Li, Kaining, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
di: Sharma, Shivam, et al.
Pubblicazione: (2026) -
BadScan: An Architectural Backdoor Attack on Visual State Space Models
di: Deshmukh, Om Suhas, et al.
Pubblicazione: (2024) -
TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering
di: Rajkhowa, Tonmoy, et al.
Pubblicazione: (2024) -
PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language
di: Haq, Ijazul, et al.
Pubblicazione: (2025) -
Invizo: Arabic Handwritten Document Optical Character Recognition Solution
di: Waly, Alhossien, et al.
Pubblicazione: (2025)