Do LLMs Have Visualization Literacy? An Evaluation on Modified Visualizations to Test Generalization in Data Interpretation
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Jiayi, Seto, Christian, Fan, Arlen, Maciejewski, Ross |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
by: Ali, Wajid, et al.
Published: (2025)
by: Ali, Wajid, et al.
Published: (2025)
When Does the Gittins Policy Have Asymptotically Optimal Response Time Tail?
by: Scully, Ziv, et al.
Published: (2021)
by: Scully, Ziv, et al.
Published: (2021)
Tests de ejecución continua: Integrated Visual and Auditory Continuous Performance Test (IVA/CPT) y TDAH. Una revisión
by: Susana Meneres-Sancho
Published: (2015)
by: Susana Meneres-Sancho
Published: (2015)
Achieving Consistent and Comparable CPU Evaluation
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
UPMEM Unleashed: Software Secrets for Speed
by: Chmielewski, Krystian, et al.
Published: (2025)
by: Chmielewski, Krystian, et al.
Published: (2025)
Opal: A Modular Framework for Optimizing Performance using Analytics and LLMs
by: Zaeed, Mohammad, et al.
Published: (2025)
by: Zaeed, Mohammad, et al.
Published: (2025)
How to Do Statistical Evaluations in ECE/CS Papers: A Practical Playbook for Defensible Results
by: Krishnamachari, Bhaskar
Published: (2026)
by: Krishnamachari, Bhaskar
Published: (2026)
Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Leveraging LLMs for Structured Information Extraction and Analysis from Cloud Incident Reports (Work In Progress Paper)
by: Chu, Xiaoyu, et al.
Published: (2026)
by: Chu, Xiaoyu, et al.
Published: (2026)
Testing the Unknown: A Framework for OpenMP Testing via Random Program Generation
by: Laguna, Ignacio, et al.
Published: (2024)
by: Laguna, Ignacio, et al.
Published: (2024)
Performance Evaluation of Subroutines Call in PHP
by: Kalmukov, Yordan
Published: (2026)
by: Kalmukov, Yordan
Published: (2026)
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
by: Huang, Haochen, et al.
Published: (2025)
by: Huang, Haochen, et al.
Published: (2025)
Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data Loaders
by: Iglovikov, Vladimir, et al.
Published: (2026)
by: Iglovikov, Vladimir, et al.
Published: (2026)
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
by: Holmes, Connor, et al.
Published: (2024)
by: Holmes, Connor, et al.
Published: (2024)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
by: Kong, Linghao, et al.
Published: (2026)
by: Kong, Linghao, et al.
Published: (2026)
Evaluation of Domain-Specific Architectures for General-Purpose Applications in Apple Silicon
by: López, Álvaro Corrochano, et al.
Published: (2025)
by: López, Álvaro Corrochano, et al.
Published: (2025)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
by: Bogart, Christopher, et al.
Published: (2025)
by: Bogart, Christopher, et al.
Published: (2025)
SynthEval: A Framework for Detailed Utility and Privacy Evaluation of Tabular Synthetic Data
by: Lautrup, Anton Danholt, et al.
Published: (2024)
by: Lautrup, Anton Danholt, et al.
Published: (2024)
Performance Evaluation of CMOS Annealing with Support Vector Machine
by: Fukuhara, Ryoga, et al.
Published: (2024)
by: Fukuhara, Ryoga, et al.
Published: (2024)
Skeptik: A Hybrid Framework for Combating Potential Misinformation in Journalism
by: Fan, Arlen, et al.
Published: (2025)
by: Fan, Arlen, et al.
Published: (2025)
Industry 4.0 Connectors -- A Performance Experiment with Modbus/TCP
by: Nikolajew, Christian, et al.
Published: (2024)
by: Nikolajew, Christian, et al.
Published: (2024)
Evaluating the Performance of the DeepSeek Model in Confidential Computing Environment
by: Dong, Ben, et al.
Published: (2025)
by: Dong, Ben, et al.
Published: (2025)
Streaming Data in HPC Workflows Using ADIOS
by: Eisenhauer, Greg, et al.
Published: (2024)
by: Eisenhauer, Greg, et al.
Published: (2024)
A Microbenchmark Framework for Performance Evaluation of OpenMP Target Offloading
by: Atif, Mohammad, et al.
Published: (2025)
by: Atif, Mohammad, et al.
Published: (2025)
Systematic Performance Evaluation Framework for LEO Mega-Constellation Satellite Networks
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
The Multiserver-Job Stochastic Recurrence Equation for Cloud Computing Performance Evaluation
by: Baccelli, Francois, et al.
Published: (2026)
by: Baccelli, Francois, et al.
Published: (2026)
Efficient Data-Driven Production Scheduling in Pharmaceutical Manufacturing
by: Balatsos, Ioannis, et al.
Published: (2026)
by: Balatsos, Ioannis, et al.
Published: (2026)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Swarm: Co-Activation Aware KVCache Offloading Across Multiple SSDs
by: Wang, Tuowei, et al.
Published: (2026)
by: Wang, Tuowei, et al.
Published: (2026)
Analysis of Stable Vertex Values: Fast Query Evaluation Over An Evolving Graph
by: Afarin, Mahbod, et al.
Published: (2025)
by: Afarin, Mahbod, et al.
Published: (2025)
A Dataset of Performance Measurements and Alerts from Mozilla (Data Artifact)
by: Besbes, Mohamed Bilel, et al.
Published: (2025)
by: Besbes, Mohamed Bilel, et al.
Published: (2025)
Visual Insights into Agentic Optimization of Pervasive Stream Processing Services
by: Sedlak, Boris, et al.
Published: (2026)
by: Sedlak, Boris, et al.
Published: (2026)
Evaluating the impact of the L3 cache size of AMD EPYC CPUs on the performance of CFD applications
by: Lawenda, Marcin, et al.
Published: (2025)
by: Lawenda, Marcin, et al.
Published: (2025)
Input-Gen: Guided Generation of Stateful Inputs for Testing, Tuning, and Training
by: Ivanov, Ivan R., et al.
Published: (2024)
by: Ivanov, Ivan R., et al.
Published: (2024)
On the Benefits of Traffic "Reprofiling'' -- The Single Hop Case
by: Song, Jiayi, et al.
Published: (2021)
by: Song, Jiayi, et al.
Published: (2021)
Integrating High Performance In-Memory Data Streaming and In-Situ Visualization in Hybrid MPI+OpenMP PIC MC Simulations Towards Exascale
by: Williams, Jeremy J., et al.
Published: (2025)
by: Williams, Jeremy J., et al.
Published: (2025)
HPL-MxP Benchmark: Mixed-Precision Algorithms, Iterative Refinement, and Scalable Data Generation
by: Dongarra, Jack, et al.
Published: (2025)
by: Dongarra, Jack, et al.
Published: (2025)
HistogramTools for Efficient Data Analysis and Distribution Representation in Large Data Sets
by: Malhotra, Shubham
Published: (2025)
by: Malhotra, Shubham
Published: (2025)
Heterogeneous Data Access Model for Concurrency Control and Methods to Deal with High Data Contention
by: Thomasian, Alexander
Published: (2024)
by: Thomasian, Alexander
Published: (2024)
Scalable Packed Layouts for Vector-Length-Agnostic ML Code Generation
by: Beysel, Ege, et al.
Published: (2026)
by: Beysel, Ege, et al.
Published: (2026)
Similar Items
-
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
by: Ali, Wajid, et al.
Published: (2025) -
When Does the Gittins Policy Have Asymptotically Optimal Response Time Tail?
by: Scully, Ziv, et al.
Published: (2021) -
Tests de ejecución continua: Integrated Visual and Auditory Continuous Performance Test (IVA/CPT) y TDAH. Una revisión
by: Susana Meneres-Sancho
Published: (2015) -
Achieving Consistent and Comparable CPU Evaluation
by: Wang, Chenxi, et al.
Published: (2024) -
UPMEM Unleashed: Software Secrets for Speed
by: Chmielewski, Krystian, et al.
Published: (2025)