Can Visual Language Models Replace OCR-Based Visual Question Answering Pipelines in Production? A Case Study in Retail
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lamm, Bianca, Keuper, Janis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Visual RAG Pipeline for Few-Shot Fine-Grained Product Classification
von: Lamm, Bianca, et al.
Veröffentlicht: (2025)
von: Lamm, Bianca, et al.
Veröffentlicht: (2025)
Retail-786k: a Large-Scale Dataset for Visual Entity Matching
von: Lamm, Bianca, et al.
Veröffentlicht: (2023)
von: Lamm, Bianca, et al.
Veröffentlicht: (2023)
Deepfakes: we need to re-think the concept of "real" images
von: Keuper, Janis, et al.
Veröffentlicht: (2025)
von: Keuper, Janis, et al.
Veröffentlicht: (2025)
Can Biases in ImageNet Models Explain Generalization?
von: Gavrikov, Paul, et al.
Veröffentlicht: (2024)
von: Gavrikov, Paul, et al.
Veröffentlicht: (2024)
PhysicsGen: Can Generative Models Learn from Images to Predict Complex Physical Relations?
von: Spitznagel, Martin, et al.
Veröffentlicht: (2025)
von: Spitznagel, Martin, et al.
Veröffentlicht: (2025)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
von: Pintore, Marco, et al.
Veröffentlicht: (2025)
As large as it gets: Learning infinitely large Filters via Neural Implicit Functions in the Fourier Domain
von: Grabinski, Julia, et al.
Veröffentlicht: (2023)
von: Grabinski, Julia, et al.
Veröffentlicht: (2023)
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
von: Shen, Zhixuan, et al.
Veröffentlicht: (2024)
von: Shen, Zhixuan, et al.
Veröffentlicht: (2024)
A New Kind of Network? Review and Reference Implementation of Neural Cellular Automata
von: Spitznagel, Martin, et al.
Veröffentlicht: (2026)
von: Spitznagel, Martin, et al.
Veröffentlicht: (2026)
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
von: Shah, Monika, et al.
Veröffentlicht: (2025)
von: Shah, Monika, et al.
Veröffentlicht: (2025)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
Efficient Visual Question Answering Pipeline for Autonomous Driving via Scene Region Compression
von: Cai, Yuliang, et al.
Veröffentlicht: (2026)
von: Cai, Yuliang, et al.
Veröffentlicht: (2026)
Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
Unfolding Local Growth Rate Estimates for (Almost) Perfect Adversarial Detection
von: Lorenz, Peter, et al.
Veröffentlicht: (2022)
von: Lorenz, Peter, et al.
Veröffentlicht: (2022)
Targeted Visual Prompting for Medical Visual Question Answering
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
Visually Interpretable Subtask Reasoning for Visual Question Answering
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
Assessing Foundation Models for Mold Colony Detection with Limited Training Data
von: Pichler, Henrik, et al.
Veröffentlicht: (2025)
von: Pichler, Henrik, et al.
Veröffentlicht: (2025)
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
Selectively Answering Visual Questions
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks
von: Lee, Jusung, et al.
Veröffentlicht: (2024)
von: Lee, Jusung, et al.
Veröffentlicht: (2024)
Large Vision-Language Models for Remote Sensing Visual Question Answering
von: Siripong, Surasakdi, et al.
Veröffentlicht: (2024)
von: Siripong, Surasakdi, et al.
Veröffentlicht: (2024)
Reconstruction as a Bridge for Event-Based Visual Question Answering
von: Lou, Hanyue, et al.
Veröffentlicht: (2025)
von: Lou, Hanyue, et al.
Veröffentlicht: (2025)
Questioning the Stability of Visual Question Answering
von: Rosenfeld, Amir, et al.
Veröffentlicht: (2025)
von: Rosenfeld, Amir, et al.
Veröffentlicht: (2025)
Benchmarking OCR Pipelines with Adaptive Enhancement for Multi-Domain Retail Bill Digitization
von: Gaikwad, Vijaysinh
Veröffentlicht: (2026)
von: Gaikwad, Vijaysinh
Veröffentlicht: (2026)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
von: Kim, Hongyeob, et al.
Veröffentlicht: (2025)
von: Kim, Hongyeob, et al.
Veröffentlicht: (2025)
Multimodal Rationales for Explainable Visual Question Answering
von: Li, Kun, et al.
Veröffentlicht: (2024)
von: Li, Kun, et al.
Veröffentlicht: (2024)
Evaluating Variance in Visual Question Answering Benchmarks
von: SR, Nikitha
Veröffentlicht: (2025)
von: SR, Nikitha
Veröffentlicht: (2025)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
von: Lim, Su Hyeon, et al.
Veröffentlicht: (2024)
von: Lim, Su Hyeon, et al.
Veröffentlicht: (2024)
Urban Sound Propagation: a Benchmark for 1-Step Generative Modeling of Complex Physical Systems
von: Spitznagel, Martin, et al.
Veröffentlicht: (2024)
von: Spitznagel, Martin, et al.
Veröffentlicht: (2024)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
Fusion of Domain-Adapted Vision and Language Models for Medical Visual Question Answering
von: Ha, Cuong Nhat, et al.
Veröffentlicht: (2024)
von: Ha, Cuong Nhat, et al.
Veröffentlicht: (2024)
Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering
von: Chen, Zhuohong, et al.
Veröffentlicht: (2026)
von: Chen, Zhuohong, et al.
Veröffentlicht: (2026)
Adapting Lightweight Vision Language Models for Radiological Visual Question Answering
von: Shourya, Aditya, et al.
Veröffentlicht: (2025)
von: Shourya, Aditya, et al.
Veröffentlicht: (2025)
Is RobustBench/AutoAttack a suitable Benchmark for Adversarial Robustness?
von: Lorenz, Peter, et al.
Veröffentlicht: (2021)
von: Lorenz, Peter, et al.
Veröffentlicht: (2021)
Adversarial Examples are Misaligned in Diffusion Model Manifolds
von: Lorenz, Peter, et al.
Veröffentlicht: (2024)
von: Lorenz, Peter, et al.
Veröffentlicht: (2024)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
von: Romero, David, et al.
Veröffentlicht: (2024)
von: Romero, David, et al.
Veröffentlicht: (2024)
How Do Training Methods Influence the Utilization of Vision Models?
von: Gavrikov, Paul, et al.
Veröffentlicht: (2024)
von: Gavrikov, Paul, et al.
Veröffentlicht: (2024)
VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Visual RAG Pipeline for Few-Shot Fine-Grained Product Classification
von: Lamm, Bianca, et al.
Veröffentlicht: (2025) -
Retail-786k: a Large-Scale Dataset for Visual Entity Matching
von: Lamm, Bianca, et al.
Veröffentlicht: (2023) -
Deepfakes: we need to re-think the concept of "real" images
von: Keuper, Janis, et al.
Veröffentlicht: (2025) -
Can Biases in ImageNet Models Explain Generalization?
von: Gavrikov, Paul, et al.
Veröffentlicht: (2024) -
PhysicsGen: Can Generative Models Learn from Images to Predict Complex Physical Relations?
von: Spitznagel, Martin, et al.
Veröffentlicht: (2025)