Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chandhok, Shivam, Fan, Wan-Cyuan, Sigal, Leonid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
von: Chandhok, Shivam, et al.
Veröffentlicht: (2025)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2025)
Test-Time Consistency in Vision Language Models
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025)
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024)
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
Tinted Frames: Question Framing Blinds Vision-Language Models
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2026)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2026)
SceneGPT: A Language Model for 3D Scene Understanding
von: Chandhok, Shivam
Veröffentlicht: (2024)
von: Chandhok, Shivam
Veröffentlicht: (2024)
Do Vision-Language Foundational models show Robust Visual Perception?
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
von: Goyal, Raghav, et al.
Veröffentlicht: (2023)
von: Goyal, Raghav, et al.
Veröffentlicht: (2023)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
von: Luo, Jiayun, et al.
Veröffentlicht: (2025)
von: Luo, Jiayun, et al.
Veröffentlicht: (2025)
On Pre-training of Multimodal Language Models Customized for Chart Understanding
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024)
Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
von: Chandhok, Shivam, et al.
Veröffentlicht: (2025)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2025)
In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
von: He, Xiangteng, et al.
Veröffentlicht: (2025)
von: He, Xiangteng, et al.
Veröffentlicht: (2025)
SPIKE-RL: Video-LLMs meet Bayesian Surprise
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
von: Salamatian, Ali, et al.
Veröffentlicht: (2025)
von: Salamatian, Ali, et al.
Veröffentlicht: (2025)
Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models
von: Luo, Jiayun, et al.
Veröffentlicht: (2023)
von: Luo, Jiayun, et al.
Veröffentlicht: (2023)
SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2025)
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2025)
Factorized Video Autoencoders for Efficient Generative Modelling
von: Suhail, Mohammed, et al.
Veröffentlicht: (2024)
von: Suhail, Mohammed, et al.
Veröffentlicht: (2024)
Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
von: Xu, Bicheng, et al.
Veröffentlicht: (2024)
von: Xu, Bicheng, et al.
Veröffentlicht: (2024)
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
von: Sakai, Shunsuke, et al.
Veröffentlicht: (2025)
von: Sakai, Shunsuke, et al.
Veröffentlicht: (2025)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2023)
3VL: Using Trees to Improve Vision-Language Models' Interpretability
von: Yellinek, Nir, et al.
Veröffentlicht: (2023)
von: Yellinek, Nir, et al.
Veröffentlicht: (2023)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
von: Bhatt, Gaurav, et al.
Veröffentlicht: (2024)
von: Bhatt, Gaurav, et al.
Veröffentlicht: (2024)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
von: Rahman, Tanzila, et al.
Veröffentlicht: (2026)
von: Rahman, Tanzila, et al.
Veröffentlicht: (2026)
Eyes Will Shut: A Vision-Based Next GPS Location Prediction Model by Reinforcement Learning from Visual Map Feed Back
von: Zhang, Ruixing, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixing, et al.
Veröffentlicht: (2025)
Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
von: Sharma, Shivam, et al.
Veröffentlicht: (2026)
von: Sharma, Shivam, et al.
Veröffentlicht: (2026)
CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models
von: An, Xiao, et al.
Veröffentlicht: (2024)
von: An, Xiao, et al.
Veröffentlicht: (2024)
Prompt2Perturb (P2P): Text-Guided Diffusion-Based Adversarial Attacks on Breast Ultrasound Images
von: Medghalchi, Yasamin, et al.
Veröffentlicht: (2024)
von: Medghalchi, Yasamin, et al.
Veröffentlicht: (2024)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2026)
von: Mahdizadeh, Ailar, et al.
Veröffentlicht: (2026)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
von: Rahman, Tanzila, et al.
Veröffentlicht: (2024)
von: Rahman, Tanzila, et al.
Veröffentlicht: (2024)
Exploring the Zero-Shot Capabilities of Vision-Language Models for Improving Gaze Following
von: Gupta, Anshul, et al.
Veröffentlicht: (2024)
von: Gupta, Anshul, et al.
Veröffentlicht: (2024)
Assessing the Geolocation Capabilities, Limitations and Societal Risks of Generative Vision-Language Models
von: Grainge, Oliver, et al.
Veröffentlicht: (2025)
von: Grainge, Oliver, et al.
Veröffentlicht: (2025)
Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection
von: Yu, Peipeng, et al.
Veröffentlicht: (2025)
von: Yu, Peipeng, et al.
Veröffentlicht: (2025)
Unleashing the Capabilities of Large Vision-Language Models for Intelligent Perception of Roadside Infrastructure
von: Fu, Luxuan, et al.
Veröffentlicht: (2026)
von: Fu, Luxuan, et al.
Veröffentlicht: (2026)
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
von: Vice, Jordan, et al.
Veröffentlicht: (2024)
von: Vice, Jordan, et al.
Veröffentlicht: (2024)
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2024)
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2024)
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2025)
von: Hossain, Mir Rayat Imtiaz, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
von: Chandhok, Shivam, et al.
Veröffentlicht: (2025) -
Test-Time Consistency in Vision Language Models
von: Chou, Shih-Han, et al.
Veröffentlicht: (2025) -
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
von: Chou, Shih-Han, et al.
Veröffentlicht: (2024) -
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2024) -
Tinted Frames: Question Framing Blinds Vision-Language Models
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2026)