VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
Fuente:
arXiv
Saved in:
| Main Authors: | Bandraupalli, Srihari, Purwar, Anupam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025)
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025)
M-PACE: Mother Child Framework for Multimodal Compliance
by: Verma, Shreyash, et al.
Published: (2025)
by: Verma, Shreyash, et al.
Published: (2025)
Are VLMs Really Blind
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
by: Rose, Daniel, et al.
Published: (2023)
by: Rose, Daniel, et al.
Published: (2023)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
by: Pani, Anupam, et al.
Published: (2025)
by: Pani, Anupam, et al.
Published: (2025)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
by: Mayer, Julius, et al.
Published: (2025)
by: Mayer, Julius, et al.
Published: (2025)
On the Perception Bottleneck of VLMs for Chart Understanding
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
CIVET: Systematic Evaluation of Understanding in VLMs
by: Rizzoli, Massimo, et al.
Published: (2025)
by: Rizzoli, Massimo, et al.
Published: (2025)
Bridging the Gap Between Multimodal Foundation Models and World Models
by: He, Xuehai
Published: (2025)
by: He, Xuehai
Published: (2025)
[De|Re]constructing VLMs' Reasoning in Counting
by: Alghisi, Simone, et al.
Published: (2025)
by: Alghisi, Simone, et al.
Published: (2025)
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
by: Stogiannidis, Ilias, et al.
Published: (2025)
by: Stogiannidis, Ilias, et al.
Published: (2025)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
by: Huang, Jen-Tse, et al.
Published: (2025)
by: Huang, Jen-Tse, et al.
Published: (2025)
Gaze-Regularized VLMs for Ego-Centric Behavior Understanding
by: Pani, Anupam, et al.
Published: (2026)
by: Pani, Anupam, et al.
Published: (2026)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
by: Yang, Cheng, et al.
Published: (2026)
by: Yang, Cheng, et al.
Published: (2026)
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
by: Bhaskar, Paramananda, et al.
Published: (2026)
by: Bhaskar, Paramananda, et al.
Published: (2026)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
by: Sun, Kaiser, et al.
Published: (2026)
by: Sun, Kaiser, et al.
Published: (2026)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
Reconstructing Animals and the Wild
by: Kulits, Peter, et al.
Published: (2024)
by: Kulits, Peter, et al.
Published: (2024)
Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
by: Poggi, Nicolas, et al.
Published: (2025)
by: Poggi, Nicolas, et al.
Published: (2025)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
by: Shi, Chufan, et al.
Published: (2026)
by: Shi, Chufan, et al.
Published: (2026)
Level Up Your Tutorials: VLMs for Game Tutorials Quality Assessment
by: Cambrin, Daniele Rege, et al.
Published: (2024)
by: Cambrin, Daniele Rege, et al.
Published: (2024)
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
by: He, Jixuan, et al.
Published: (2026)
by: He, Jixuan, et al.
Published: (2026)
PromptSync: Bridging Domain Gaps in Vision-Language Models through Class-Aware Prototype Alignment and Discrimination
by: Khandelwal, Anant
Published: (2024)
by: Khandelwal, Anant
Published: (2024)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
Multimodal Event Detection: Current Approaches and Defining the New Playground through LLMs and VLMs
by: Dey, Abhishek, et al.
Published: (2025)
by: Dey, Abhishek, et al.
Published: (2025)
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
by: Ye, Junyi, et al.
Published: (2024)
by: Ye, Junyi, et al.
Published: (2024)
BridgeTower: Building Bridges Between Encoders in Vision-Language Representation Learning
by: Xu, Xiao, et al.
Published: (2022)
by: Xu, Xiao, et al.
Published: (2022)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
by: Yang, Qian, et al.
Published: (2024)
by: Yang, Qian, et al.
Published: (2024)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
by: Shahgir, Haz Sameen, et al.
Published: (2026)
by: Shahgir, Haz Sameen, et al.
Published: (2026)
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies
by: Gao, Yingqiang, et al.
Published: (2024)
by: Gao, Yingqiang, et al.
Published: (2024)
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
by: Lee, DongGeon, et al.
Published: (2025)
by: Lee, DongGeon, et al.
Published: (2025)
Real-world Instance-specific Image Goal Navigation: Bridging Domain Gaps via Contrastive Learning
by: Sakaguchi, Taichi, et al.
Published: (2024)
by: Sakaguchi, Taichi, et al.
Published: (2024)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024)
by: Yanuka, Moran, et al.
Published: (2024)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
by: Tao, Wei, et al.
Published: (2026)
by: Tao, Wei, et al.
Published: (2026)
Leveraging NTPs for Efficient Hallucination Detection in VLMs
by: Azachi, Ofir, et al.
Published: (2025)
by: Azachi, Ofir, et al.
Published: (2025)
Understanding and Rectifying Safety Perception Distortion in VLMs
by: Zou, Xiaohan, et al.
Published: (2025)
by: Zou, Xiaohan, et al.
Published: (2025)
Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization
by: Islam, Md Moinul, et al.
Published: (2025)
by: Islam, Md Moinul, et al.
Published: (2025)
Similar Items
-
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
by: Liu, Peng, et al.
Published: (2025) -
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025) -
M-PACE: Mother Child Framework for Multimodal Compliance
by: Verma, Shreyash, et al.
Published: (2025) -
Are VLMs Really Blind
by: Singh, Ayush, et al.
Published: (2024) -
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
by: Rose, Daniel, et al.
Published: (2023)