Exploring MLLMs Perception of Network Visualization Principles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miller, Jacob, Wallinger, Markus, Felder, Ludwig, Brand, Timo, Förster, Henry, Zink, Johannes, Chen, Chunyang, Kobourov, Stephen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913017143230464
author Miller, Jacob
Wallinger, Markus
Felder, Ludwig
Brand, Timo
Förster, Henry
Zink, Johannes
Chen, Chunyang
Kobourov, Stephen
author_facet Miller, Jacob
Wallinger, Markus
Felder, Ludwig
Brand, Timo
Förster, Henry
Zink, Johannes
Chen, Chunyang
Kobourov, Stephen
contents In this paper, we test whether Multimodal Large Language Models (MLLMs) can match human-subject performance in tasks involving the perception of properties in network layouts. Specifically, we replicate a human-subject experiment about perceiving quality (namely stress) in network layouts using GPT-4o, Gemini-2.5 and Qwen2.5. Our experiments show that giving MLLMs the same study information as trained human participants yields performance comparable to that of human experts and exceeds that of untrained non-experts. Additionally, we show that prompt engineering that deviates from the human-subject experiment can lead to better-than-human performance in some settings. Interestingly, like human subjects, the MLLMs seem to rely on visual proxies rather than computing the actual value of stress, indicating some sense or facsimile of perception. Explanations from the models are similar to those used by the human participants (e.g., an even distribution of nodes and uniform edge lengths).
format Preprint
id arxiv_https___arxiv_org_abs_2506_14611
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring MLLMs Perception of Network Visualization Principles
Miller, Jacob
Wallinger, Markus
Felder, Ludwig
Brand, Timo
Förster, Henry
Zink, Johannes
Chen, Chunyang
Kobourov, Stephen
Human-Computer Interaction
In this paper, we test whether Multimodal Large Language Models (MLLMs) can match human-subject performance in tasks involving the perception of properties in network layouts. Specifically, we replicate a human-subject experiment about perceiving quality (namely stress) in network layouts using GPT-4o, Gemini-2.5 and Qwen2.5. Our experiments show that giving MLLMs the same study information as trained human participants yields performance comparable to that of human experts and exceeds that of untrained non-experts. Additionally, we show that prompt engineering that deviates from the human-subject experiment can lead to better-than-human performance in some settings. Interestingly, like human subjects, the MLLMs seem to rely on visual proxies rather than computing the actual value of stress, indicating some sense or facsimile of perception. Explanations from the models are similar to those used by the human participants (e.g., an even distribution of nodes and uniform edge lengths).
title Exploring MLLMs Perception of Network Visualization Principles
topic Human-Computer Interaction
url https://arxiv.org/abs/2506.14611