Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Grace, Kang, Jenna Jiayi, Shah, Raj Sanjay, Pfister, Hanspeter, Varma, Sashank |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations
by: Lalai, Harsh Nishant, et al.
Published: (2026)
by: Lalai, Harsh Nishant, et al.
Published: (2026)
Tree of Attributes Prompt Learning for Vision-Language Models
by: Ding, Tong, et al.
Published: (2024)
by: Ding, Tong, et al.
Published: (2024)
Computer Vision Modeling of the Development of Geometric and Numerical Concepts in Humans
by: Wang, Zekun, et al.
Published: (2025)
by: Wang, Zekun, et al.
Published: (2025)
Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts
by: Wang, Zekun, et al.
Published: (2025)
by: Wang, Zekun, et al.
Published: (2025)
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories
by: Shah, Raj Sanjay, et al.
Published: (2025)
by: Shah, Raj Sanjay, et al.
Published: (2025)
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
by: Xu, Tianling, et al.
Published: (2025)
by: Xu, Tianling, et al.
Published: (2025)
Prompt-based Visual Alignment for Zero-shot Policy Transfer
by: Gao, Haihan, et al.
Published: (2024)
by: Gao, Haihan, et al.
Published: (2024)
Development of Cognitive Intelligence in Pre-trained Language Models
by: Shah, Raj Sanjay, et al.
Published: (2024)
by: Shah, Raj Sanjay, et al.
Published: (2024)
CTRL-GS: Cascaded Temporal Residue Learning for 4D Gaussian Splatting
by: Hou, Karly, et al.
Published: (2025)
by: Hou, Karly, et al.
Published: (2025)
DualEdit: Dual Editing for Knowledge Updating in Vision-Language Models
by: Shi, Zhiyi, et al.
Published: (2025)
by: Shi, Zhiyi, et al.
Published: (2025)
SocialGPT: Prompting LLMs for Social Relation Reasoning via Greedy Segment Optimization
by: Li, Wanhua, et al.
Published: (2024)
by: Li, Wanhua, et al.
Published: (2024)
LangFlash: Feed-forward 3D Language Gaussian Splatting from Sparse Unposed Images
by: Liu, Yilong, et al.
Published: (2026)
by: Liu, Yilong, et al.
Published: (2026)
Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
by: Tang, Lv, et al.
Published: (2023)
by: Tang, Lv, et al.
Published: (2023)
LangSplat: 3D Language Gaussian Splatting
by: Qin, Minghan, et al.
Published: (2023)
by: Qin, Minghan, et al.
Published: (2023)
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
by: Yang, Chanhyeong, et al.
Published: (2025)
by: Yang, Chanhyeong, et al.
Published: (2025)
Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models
by: Magid, Salma Abdel, et al.
Published: (2024)
by: Magid, Salma Abdel, et al.
Published: (2024)
CoVis: A Collaborative Framework for Fine-grained Graphic Visual Understanding
by: Deng, Xiaoyu, et al.
Published: (2024)
by: Deng, Xiaoyu, et al.
Published: (2024)
GenCLIP: Generalizing CLIP Prompts for Zero-shot Anomaly Detection
by: Kim, Donghyeong, et al.
Published: (2025)
by: Kim, Donghyeong, et al.
Published: (2025)
Evaluating Graphical Perception Capabilities of Vision Transformers
by: Poonam, Poonam, et al.
Published: (2026)
by: Poonam, Poonam, et al.
Published: (2026)
VGDiffZero: Text-to-image Diffusion Models Can Be Zero-shot Visual Grounders
by: Liu, Xuyang, et al.
Published: (2023)
by: Liu, Xuyang, et al.
Published: (2023)
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
by: Shen, Yang, et al.
Published: (2024)
by: Shen, Yang, et al.
Published: (2024)
Vision and Language Reference Prompt into SAM for Few-shot Segmentation
by: Sakurai, Kosuke, et al.
Published: (2025)
by: Sakurai, Kosuke, et al.
Published: (2025)
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
by: Wang, Zhecan, et al.
Published: (2023)
by: Wang, Zhecan, et al.
Published: (2023)
Differentiable Inverse Graphics for Zero-shot Scene Reconstruction and Robot Grasping
by: Arriaga, Octavio, et al.
Published: (2026)
by: Arriaga, Octavio, et al.
Published: (2026)
Bias at the End of the Score
by: Magid, Salma Abdel, et al.
Published: (2026)
by: Magid, Salma Abdel, et al.
Published: (2026)
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding
by: Sun, Qi, et al.
Published: (2024)
by: Sun, Qi, et al.
Published: (2024)
Pre-training LLMs using human-like development data corpus
by: Bhardwaj, Khushi, et al.
Published: (2023)
by: Bhardwaj, Khushi, et al.
Published: (2023)
View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
by: Liu, Yuanyuan, et al.
Published: (2025)
by: Liu, Yuanyuan, et al.
Published: (2025)
Zero-shot Generalizable Incremental Learning for Vision-Language Object Detection
by: Deng, Jieren, et al.
Published: (2024)
by: Deng, Jieren, et al.
Published: (2024)
AttriPrompter: Auto-Prompting with Attribute Semantics for Zero-shot Nuclei Detection via Visual-Language Pre-trained Models
by: Wu, Yongjian, et al.
Published: (2024)
by: Wu, Yongjian, et al.
Published: (2024)
RiGS: Rigid-aware 4D Gaussian Splatting from a Single Monocular Video
by: Wu, Chenyu, et al.
Published: (2026)
by: Wu, Chenyu, et al.
Published: (2026)
Zero-shot Vision-Language Reranking for Cross-View Geolocalization
by: Erzurumlu, Yunus Talha, et al.
Published: (2026)
by: Erzurumlu, Yunus Talha, et al.
Published: (2026)
Label Propagation for Zero-shot Classification with Vision-Language Models
by: Stojnić, Vladan, et al.
Published: (2024)
by: Stojnić, Vladan, et al.
Published: (2024)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
by: Aklilu, Josiah, et al.
Published: (2024)
by: Aklilu, Josiah, et al.
Published: (2024)
TriSAM: Tri-Plane SAM for zero-shot cortical blood vessel segmentation in VEM images
by: Wan, Jia, et al.
Published: (2024)
by: Wan, Jia, et al.
Published: (2024)
Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography
by: Spinaci, Gianmarco, et al.
Published: (2025)
by: Spinaci, Gianmarco, et al.
Published: (2025)
Causal Graphical Models for Vision-Language Compositional Understanding
by: Parascandolo, Fiorenzo, et al.
Published: (2024)
by: Parascandolo, Fiorenzo, et al.
Published: (2024)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
by: Ma, Shuailei, et al.
Published: (2023)
by: Ma, Shuailei, et al.
Published: (2023)
Similar Items
-
When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations
by: Lalai, Harsh Nishant, et al.
Published: (2026) -
Tree of Attributes Prompt Learning for Vision-Language Models
by: Ding, Tong, et al.
Published: (2024) -
Computer Vision Modeling of the Development of Geometric and Numerical Concepts in Humans
by: Wang, Zekun, et al.
Published: (2025) -
Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts
by: Wang, Zekun, et al.
Published: (2025) -
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
by: Liu, Yifan, et al.
Published: (2025)