Benchmarking and Evaluating VLMs for Software Architecture Diagram Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Ouyang, Shuyin, Zhang, Jie M., Gong, Jingzhi, Jahangirova, Gunel, Mousavi, Mohammad Reza, Johns, Jack, Lee, Beum Seuk, Ziolkowski, Adam, Virginas, Botond, Noppen, Joost |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Taxonomy of Real Faults in Hybrid Quantum-Classical Architectures
by: Bensoussan, Avner, et al.
Published: (2025)
by: Bensoussan, Avner, et al.
Published: (2025)
The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project
by: Gröpler, Robin, et al.
Published: (2025)
by: Gröpler, Robin, et al.
Published: (2025)
How Does Chunking Affect Retrieval-Augmented Code Completion? A Controlled Empirical Study
by: Wu, Xinjian, et al.
Published: (2026)
by: Wu, Xinjian, et al.
Published: (2026)
Understanding LLM-Driven Test Oracle Generation
by: Bodicoat, Adam, et al.
Published: (2026)
by: Bodicoat, Adam, et al.
Published: (2026)
Comparative Analysis of Carbon Footprint in Manual vs. LLM-Assisted Code Development
by: Cheung, Kuen Sum, et al.
Published: (2025)
by: Cheung, Kuen Sum, et al.
Published: (2025)
Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
by: Jahangirova, Gunel, et al.
Published: (2024)
by: Jahangirova, Gunel, et al.
Published: (2024)
(How) Do Large Language Models Understand High-Level Message Sequence Charts?
by: Mousavi, Mohammad Reza
Published: (2026)
by: Mousavi, Mohammad Reza
Published: (2026)
MuFF: Stable and Sensitive Post-training Mutation Testing for Deep Learning
by: Kim, Jinhan, et al.
Published: (2025)
by: Kim, Jinhan, et al.
Published: (2025)
GenMorph: Automatically Generating Metamorphic Relations via Genetic Programming
by: Ayerdi, Jon, et al.
Published: (2023)
by: Ayerdi, Jon, et al.
Published: (2023)
An Empirical Study of Fault Localisation Techniques for Deep Learning
by: Humbatova, Nargiz, et al.
Published: (2024)
by: Humbatova, Nargiz, et al.
Published: (2024)
New Formulation of DNN Statistical Mutation Killing for Ensuring Monotonicity: A Technical Report
by: Kim, Jinhan, et al.
Published: (2025)
by: Kim, Jinhan, et al.
Published: (2025)
Revisiting "Revisiting Neuron Coverage for DNN Testing: A Layer-Wise and Distribution-Aware Criterion": A Critical Review and Implications on DNN Coverage Testing
by: Kim, Jinhan, et al.
Published: (2026)
by: Kim, Jinhan, et al.
Published: (2026)
Fault Localisation and Repair for DL Systems: An Empirical Study with LLMs
by: Kim, Jinhan, et al.
Published: (2025)
by: Kim, Jinhan, et al.
Published: (2025)
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
by: Ravi, Ravin, et al.
Published: (2026)
by: Ravi, Ravin, et al.
Published: (2026)
Predicting Software Performance with Divide-and-Learn
by: Gong, Jingzhi, et al.
Published: (2023)
by: Gong, Jingzhi, et al.
Published: (2023)
muPRL: A Mutation Testing Pipeline for Deep Reinforcement Learning based on Real Faults
by: Thomas, Deepak-George, et al.
Published: (2024)
by: Thomas, Deepak-George, et al.
Published: (2024)
CIVET: Systematic Evaluation of Understanding in VLMs
by: Rizzoli, Massimo, et al.
Published: (2025)
by: Rizzoli, Massimo, et al.
Published: (2025)
Towards Living Software Architecture Diagrams
by: Correia, Filipe F., et al.
Published: (2024)
by: Correia, Filipe F., et al.
Published: (2024)
Software is infrastructure: failures, successes, costs, and the case for formal verification
by: Bernardi, Giovanni, et al.
Published: (2025)
by: Bernardi, Giovanni, et al.
Published: (2025)
Enhancing Trust in Language Model-Based Code Optimization through RLHF: A Research Design
by: Gong, Jingzhi
Published: (2025)
by: Gong, Jingzhi
Published: (2025)
Pushing the Boundary: Specialising Deep Configuration Performance Learning
by: Gong, Jingzhi
Published: (2024)
by: Gong, Jingzhi
Published: (2024)
Helveg: Diagrams for Software Documentation
by: Štěpánek, Adam, et al.
Published: (2025)
by: Štěpánek, Adam, et al.
Published: (2025)
Interactive Diagrams for Software Documentation
by: Štěpánek, Adam, et al.
Published: (2024)
by: Štěpánek, Adam, et al.
Published: (2024)
Learning Software Bug Reports: A Systematic Literature Review
by: Long, Guoming, et al.
Published: (2025)
by: Long, Guoming, et al.
Published: (2025)
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
by: Xiang, Yanzheng, et al.
Published: (2025)
by: Xiang, Yanzheng, et al.
Published: (2025)
DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
by: Faraz, Ali, et al.
Published: (2025)
by: Faraz, Ali, et al.
Published: (2025)
Complete FSM Testing Using Strong Separability
by: Hierons, Robert M., et al.
Published: (2025)
by: Hierons, Robert M., et al.
Published: (2025)
Toward Live Noise Fingerprinting in Quantum Software Engineering
by: Bensoussan, Avner, et al.
Published: (2025)
by: Bensoussan, Avner, et al.
Published: (2025)
Floating Power
by: Günel, Gökçe
Published: (2026)
by: Günel, Gökçe
Published: (2026)
An impedance matching section based on composite right/left‐handed transmission line model in viewpoint of phase shift and noise figure
by: Tayfun Günel
Published: (2024)
by: Tayfun Günel
Published: (2024)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
by: Ouyang, Shuyin, et al.
Published: (2025)
by: Ouyang, Shuyin, et al.
Published: (2025)
Deep Configuration Performance Learning: A Systematic Survey and Taxonomy
by: Gong, Jingzhi, et al.
Published: (2024)
by: Gong, Jingzhi, et al.
Published: (2024)
Predicting Configuration Performance in Multiple Environments with Sequential Meta-learning
by: Gong, Jingzhi, et al.
Published: (2024)
by: Gong, Jingzhi, et al.
Published: (2024)
A Study of LLMs' Preferences for Libraries and Programming Languages
by: Twist, Lukas, et al.
Published: (2025)
by: Twist, Lukas, et al.
Published: (2025)
Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
by: Narayanan, Aravind, et al.
Published: (2025)
by: Narayanan, Aravind, et al.
Published: (2025)
Energy demand pattern analysis in South Korea using hidden Markov model‐based classification
by: Jaeyong Lee, et al.
Published: (2024)
by: Jaeyong Lee, et al.
Published: (2024)
Ordered probit Bayesian additive regression trees for ordinal data
by: Jaeyong Lee, et al.
Published: (2024)
by: Jaeyong Lee, et al.
Published: (2024)
[De|Re]constructing VLMs' Reasoning in Counting
by: Alghisi, Simone, et al.
Published: (2025)
by: Alghisi, Simone, et al.
Published: (2025)
Reading the Juggler of Notre Dame
by: Ziolkowski, Jan
Published: (2022)
by: Ziolkowski, Jan
Published: (2022)
Similar Items
-
A Taxonomy of Real Faults in Hybrid Quantum-Classical Architectures
by: Bensoussan, Avner, et al.
Published: (2025) -
The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project
by: Gröpler, Robin, et al.
Published: (2025) -
How Does Chunking Affect Retrieval-Augmented Code Completion? A Controlled Empirical Study
by: Wu, Xinjian, et al.
Published: (2026) -
Understanding LLM-Driven Test Oracle Generation
by: Bodicoat, Adam, et al.
Published: (2026) -
Comparative Analysis of Carbon Footprint in Manual vs. LLM-Assisted Code Development
by: Cheung, Kuen Sum, et al.
Published: (2025)