CANVAS: A Benchmark for Vision-Language Models on Tool-Based User Interface Design
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jeong, Daeheon, Byun, Seoyeon, Son, Kihoon, Kim, Dae Hyun, Kim, Juho |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
von: Kim, Mingyeong, et al.
Veröffentlicht: (2026)
von: Kim, Mingyeong, et al.
Veröffentlicht: (2026)
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
von: Kim, Dain, et al.
Veröffentlicht: (2026)
von: Kim, Dain, et al.
Veröffentlicht: (2026)
GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models
von: Wu, Zixuan, et al.
Veröffentlicht: (2024)
von: Wu, Zixuan, et al.
Veröffentlicht: (2024)
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
von: Kim, Minchan, et al.
Veröffentlicht: (2024)
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models
von: Ju, Jeongho, et al.
Veröffentlicht: (2024)
von: Ju, Jeongho, et al.
Veröffentlicht: (2024)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
von: Kim, Yoonshik, et al.
Veröffentlicht: (2025)
von: Kim, Yoonshik, et al.
Veröffentlicht: (2025)
Decoding fMRI Data into Captions using Prefix Language Modeling
von: Shen, Vyacheslav, et al.
Veröffentlicht: (2025)
von: Shen, Vyacheslav, et al.
Veröffentlicht: (2025)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2024)
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2024)
Intriguing Properties of Large Language and Vision Models
von: Lee, Young-Jun, et al.
Veröffentlicht: (2024)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2024)
Vision-Language Models Do Not Understand Negation
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
von: Kim, Geewook, et al.
Veröffentlicht: (2024)
von: Kim, Geewook, et al.
Veröffentlicht: (2024)
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
debiaSAE: Benchmarking and Mitigating Vision-Language Model Bias
von: Sasse, Kuleen, et al.
Veröffentlicht: (2024)
von: Sasse, Kuleen, et al.
Veröffentlicht: (2024)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
von: Hu, Yushi, et al.
Veröffentlicht: (2023)
von: Hu, Yushi, et al.
Veröffentlicht: (2023)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
von: Gröpl, Marcel, et al.
Veröffentlicht: (2026)
von: Gröpl, Marcel, et al.
Veröffentlicht: (2026)
TuneVLSeg: Prompt Tuning Benchmark for Vision-Language Segmentation Models
von: Adhikari, Rabin, et al.
Veröffentlicht: (2024)
von: Adhikari, Rabin, et al.
Veröffentlicht: (2024)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
von: He, Xingwei, et al.
Veröffentlicht: (2024)
von: He, Xingwei, et al.
Veröffentlicht: (2024)
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
von: Wu, Xiyang, et al.
Veröffentlicht: (2024)
von: Wu, Xiyang, et al.
Veröffentlicht: (2024)
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
UIClip: A Data-driven Model for Assessing User Interface Design
von: Wu, Jason, et al.
Veröffentlicht: (2024)
von: Wu, Jason, et al.
Veröffentlicht: (2024)
Doubly-Universal Adversarial Perturbations: Deceiving Vision-Language Models Across Both Images and Text with a Single Perturbation
von: Kim, Hee-Seon, et al.
Veröffentlicht: (2024)
von: Kim, Hee-Seon, et al.
Veröffentlicht: (2024)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays
von: Cho, Yeongjae, et al.
Veröffentlicht: (2024)
von: Cho, Yeongjae, et al.
Veröffentlicht: (2024)
MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024)
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
von: Chung, Jiwan, et al.
Veröffentlicht: (2024)
Bayesian Principles Improve Prompt Learning In Vision-Language Models
von: Kim, Mingyu, et al.
Veröffentlicht: (2025)
von: Kim, Mingyu, et al.
Veröffentlicht: (2025)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
von: Ray, Sourjyadip, et al.
Veröffentlicht: (2024)
von: Ray, Sourjyadip, et al.
Veröffentlicht: (2024)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
von: Son, Jaemin, et al.
Veröffentlicht: (2025)
von: Son, Jaemin, et al.
Veröffentlicht: (2025)
Benchmarking Vision Language Models for Cultural Understanding
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
von: Lee, Seongyun, et al.
Veröffentlicht: (2024) -
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
von: Kim, Mingyeong, et al.
Veröffentlicht: (2026) -
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
von: Kim, Dain, et al.
Veröffentlicht: (2026) -
GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models
von: Wu, Zixuan, et al.
Veröffentlicht: (2024) -
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)