DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kukreja, Dikshant, Sah, Kshitij, Goyal, Karan, Mohania, Mukesh, Goyal, Vikram |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAGE: Bridging the Accuracy-Aesthetics Gap in Educational Diagrams via Code-Anchored Generative Enhancement
by: Kukreja, Dikshant, et al.
Published: (2026)
by: Kukreja, Dikshant, et al.
Published: (2026)
Public Profile Matters: A Scalable Integrated Approach to Recommend Citations in the Wild
by: Goyal, Karan, et al.
Published: (2026)
by: Goyal, Karan, et al.
Published: (2026)
SymTax: Symbiotic Relationship and Taxonomy Fusion for Effective Citation Recommendation
by: Goyal, Karan, et al.
Published: (2024)
by: Goyal, Karan, et al.
Published: (2024)
The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm
by: Goyal, Karan
Published: (2026)
by: Goyal, Karan
Published: (2026)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024)
by: Li, Kevin Y., et al.
Published: (2024)
PQV-Mobile: A Combined Pruning and Quantization Toolkit to Optimize Vision Transformers for Mobile Applications
by: Bhardwaj, Kshitij
Published: (2024)
by: Bhardwaj, Kshitij
Published: (2024)
DistortBench: Benchmarking Vision Language Models on Image Distortion Identification
by: Goyal, Divyanshu, et al.
Published: (2026)
by: Goyal, Divyanshu, et al.
Published: (2026)
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
by: Hong, Hyesoo, et al.
Published: (2026)
by: Hong, Hyesoo, et al.
Published: (2026)
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
by: Shabtay, Nimrod, et al.
Published: (2026)
by: Shabtay, Nimrod, et al.
Published: (2026)
Evaluating Vision Language Models (VLMs) for Radiology: A Comprehensive Analysis
by: Li, Frank, et al.
Published: (2025)
by: Li, Frank, et al.
Published: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
by: Berman, Shmuel, et al.
Published: (2025)
by: Berman, Shmuel, et al.
Published: (2025)
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
by: Singh, Ishika, et al.
Published: (2025)
by: Singh, Ishika, et al.
Published: (2025)
CLIP-Inspector: Model-Level Backdoor Detection for Prompt-Tuned CLIP via OOD Trigger Inversion
by: Jindal, Akshit, et al.
Published: (2026)
by: Jindal, Akshit, et al.
Published: (2026)
Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
Stateful Token Reduction for Long-Video Hybrid VLMs
by: Jiang, Jindong, et al.
Published: (2026)
by: Jiang, Jindong, et al.
Published: (2026)
Towards Lossless Ultimate Vision Token Compression for VLMs
by: Zheng, Dehua, et al.
Published: (2025)
by: Zheng, Dehua, et al.
Published: (2025)
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
by: Zhang, Gengyuan, et al.
Published: (2023)
by: Zhang, Gengyuan, et al.
Published: (2023)
DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models
by: Wang, JiYang, et al.
Published: (2026)
by: Wang, JiYang, et al.
Published: (2026)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
by: Luo, Kun, et al.
Published: (2026)
by: Luo, Kun, et al.
Published: (2026)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
NanoVLMs: How small can we go and still make coherent Vision Language Models?
by: Agarwalla, Mukund, et al.
Published: (2025)
by: Agarwalla, Mukund, et al.
Published: (2025)
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
by: Tartaglini, Alexa R., et al.
Published: (2025)
by: Tartaglini, Alexa R., et al.
Published: (2025)
Edge Reliability Gap in Vision-Language Models: Quantifying Failure Modes of Compressed VLMs Under Visual Corruption
by: Erol, Mehmet Kaan
Published: (2026)
by: Erol, Mehmet Kaan
Published: (2026)
SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters
by: Shah, Arya, et al.
Published: (2026)
by: Shah, Arya, et al.
Published: (2026)
YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval
by: Pan, Jiancheng, et al.
Published: (2024)
by: Pan, Jiancheng, et al.
Published: (2024)
Revisiting the Role of Language Priors in Vision-Language Models
by: Lin, Zhiqiu, et al.
Published: (2023)
by: Lin, Zhiqiu, et al.
Published: (2023)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
by: Duan, Yicheng, et al.
Published: (2025)
by: Duan, Yicheng, et al.
Published: (2025)
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
by: Pan, Jiazhen, et al.
Published: (2025)
by: Pan, Jiazhen, et al.
Published: (2025)
Token Pruning using a Lightweight Background Aware Vision Transformer
by: Sah, Sudhakar, et al.
Published: (2024)
by: Sah, Sudhakar, et al.
Published: (2024)
Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
by: Zhong, Yuan, et al.
Published: (2025)
by: Zhong, Yuan, et al.
Published: (2025)
MapTrace: Scalable Data Generation for Route Tracing on Maps
by: Panagopoulou, Artemis, et al.
Published: (2025)
by: Panagopoulou, Artemis, et al.
Published: (2025)
Enhancing LLM-Based Neural Network Generation: Few-Shot Prompting and Efficient Validation for Automated Architecture Design
by: Duvvuri, Raghuvir, et al.
Published: (2025)
by: Duvvuri, Raghuvir, et al.
Published: (2025)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
by: Zhou, Guanyu, et al.
Published: (2026)
by: Zhou, Guanyu, et al.
Published: (2026)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
by: Lee, Kang-il, et al.
Published: (2024)
by: Lee, Kang-il, et al.
Published: (2024)
Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis
by: Samanta, Argha Kamal, et al.
Published: (2025)
by: Samanta, Argha Kamal, et al.
Published: (2025)
VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
by: Yang, Xindi, et al.
Published: (2025)
by: Yang, Xindi, et al.
Published: (2025)
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
by: Yuan, Kun, et al.
Published: (2025)
by: Yuan, Kun, et al.
Published: (2025)
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Public Computer Vision Datasets for Precision Livestock Farming: A Systematic Survey
by: Bhujel, Anil, et al.
Published: (2024)
by: Bhujel, Anil, et al.
Published: (2024)
Similar Items
-
CAGE: Bridging the Accuracy-Aesthetics Gap in Educational Diagrams via Code-Anchored Generative Enhancement
by: Kukreja, Dikshant, et al.
Published: (2026) -
Public Profile Matters: A Scalable Integrated Approach to Recommend Citations in the Wild
by: Goyal, Karan, et al.
Published: (2026) -
SymTax: Symbiotic Relationship and Taxonomy Fusion for Effective Citation Recommendation
by: Goyal, Karan, et al.
Published: (2024) -
The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm
by: Goyal, Karan
Published: (2026) -
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024)