PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Ravishankara, Mayank |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis
von: Ravishankara, Mayank
Veröffentlicht: (2026)
von: Ravishankara, Mayank
Veröffentlicht: (2026)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
von: Galimzyanov, Timur, et al.
Veröffentlicht: (2024)
von: Galimzyanov, Timur, et al.
Veröffentlicht: (2024)
Cost-Efficient Prompt Engineering for Unsupervised Entity Resolution
von: Nananukul, Navapat, et al.
Veröffentlicht: (2023)
von: Nananukul, Navapat, et al.
Veröffentlicht: (2023)
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
von: Jiang, Shufan, et al.
Veröffentlicht: (2026)
von: Jiang, Shufan, et al.
Veröffentlicht: (2026)
LLMs: A Game-Changer for Software Engineers?
von: Haque, Md Asraful
Veröffentlicht: (2024)
von: Haque, Md Asraful
Veröffentlicht: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
von: Ngassom, Sylvain Kouemo, et al.
Veröffentlicht: (2024)
von: Ngassom, Sylvain Kouemo, et al.
Veröffentlicht: (2024)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2026)
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2026)
Analysis of LLMs vs Human Experts in Requirements Engineering
von: Hymel, Cory, et al.
Veröffentlicht: (2025)
von: Hymel, Cory, et al.
Veröffentlicht: (2025)
Evaluating LLMs for Visualization Tasks
von: Khan, Saadiq Rauf, et al.
Veröffentlicht: (2025)
von: Khan, Saadiq Rauf, et al.
Veröffentlicht: (2025)
LLMs for Engineering: Teaching Models to Design High Powered Rockets
von: Simonds, Toby
Veröffentlicht: (2025)
von: Simonds, Toby
Veröffentlicht: (2025)
Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives
von: Khati, Dipin, et al.
Veröffentlicht: (2025)
von: Khati, Dipin, et al.
Veröffentlicht: (2025)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
von: Wu, Fan, et al.
Veröffentlicht: (2026)
von: Wu, Fan, et al.
Veröffentlicht: (2026)
Lifecycle-Aware code generation: Leveraging Software Engineering Phases in LLMs
von: Xing, Xing, et al.
Veröffentlicht: (2025)
von: Xing, Xing, et al.
Veröffentlicht: (2025)
Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research
von: Trinkenreich, Bianca, et al.
Veröffentlicht: (2025)
von: Trinkenreich, Bianca, et al.
Veröffentlicht: (2025)
Evaluating the Energy-Efficiency of the Code Generated by LLMs
von: Islam, Md Arman, et al.
Veröffentlicht: (2025)
von: Islam, Md Arman, et al.
Veröffentlicht: (2025)
Evaluating the Generalizability of LLMs in Automated Program Repair
von: Li, Fengjie, et al.
Veröffentlicht: (2025)
von: Li, Fengjie, et al.
Veröffentlicht: (2025)
From Inductive to Deductive: LLMs-Based Qualitative Data Analysis in Requirements Engineering
von: Shah, Syed Tauhid Ullah, et al.
Veröffentlicht: (2025)
von: Shah, Syed Tauhid Ullah, et al.
Veröffentlicht: (2025)
Holistic Evaluation of State-of-the-Art LLMs for Code Generation
von: Zhang, Le, et al.
Veröffentlicht: (2025)
von: Zhang, Le, et al.
Veröffentlicht: (2025)
Using LLMs in Software Requirements Specifications: An Empirical Evaluation
von: Krishna, Madhava, et al.
Veröffentlicht: (2024)
von: Krishna, Madhava, et al.
Veröffentlicht: (2024)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
von: Wang, Hanbin, et al.
Veröffentlicht: (2025)
von: Wang, Hanbin, et al.
Veröffentlicht: (2025)
Repository Intelligence Graph: Deterministic Architectural Map for LLM Code Assistants
von: Cherny-Shahar, Tsvi, et al.
Veröffentlicht: (2026)
von: Cherny-Shahar, Tsvi, et al.
Veröffentlicht: (2026)
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
von: Pan, Chenkai, et al.
Veröffentlicht: (2026)
von: Pan, Chenkai, et al.
Veröffentlicht: (2026)
TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance
von: Bruches, Elena, et al.
Veröffentlicht: (2026)
von: Bruches, Elena, et al.
Veröffentlicht: (2026)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
von: Singha, Ananya, et al.
Veröffentlicht: (2025)
von: Singha, Ananya, et al.
Veröffentlicht: (2025)
FullStack Bench: Evaluating LLMs as Full Stack Coders
von: Bytedance-Seed-Foundation-Code-Team, et al.
Veröffentlicht: (2024)
von: Bytedance-Seed-Foundation-Code-Team, et al.
Veröffentlicht: (2024)
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
von: Li, Jingyue, et al.
Veröffentlicht: (2026)
von: Li, Jingyue, et al.
Veröffentlicht: (2026)
ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation
von: Xianpeng, et al.
Veröffentlicht: (2026)
von: Xianpeng, et al.
Veröffentlicht: (2026)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
von: Zhu, Hongda, et al.
Veröffentlicht: (2025)
von: Zhu, Hongda, et al.
Veröffentlicht: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
von: Fandina, Ora Nova, et al.
Veröffentlicht: (2025)
von: Fandina, Ora Nova, et al.
Veröffentlicht: (2025)
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
von: Khati, Dipin, et al.
Veröffentlicht: (2026)
von: Khati, Dipin, et al.
Veröffentlicht: (2026)
An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains
von: Yan, Zihe, et al.
Veröffentlicht: (2025)
von: Yan, Zihe, et al.
Veröffentlicht: (2025)
Adaptive Hierarchical Evaluation of LLMs and SAST tools for CWE Prediction in Python
von: Adnan, Muntasir, et al.
Veröffentlicht: (2026)
von: Adnan, Muntasir, et al.
Veröffentlicht: (2026)
Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects
von: Nunes, Henrique, et al.
Veröffentlicht: (2025)
von: Nunes, Henrique, et al.
Veröffentlicht: (2025)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
von: Wang, Shufan, et al.
Veröffentlicht: (2025)
von: Wang, Shufan, et al.
Veröffentlicht: (2025)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
von: Li, Ziyu, et al.
Veröffentlicht: (2024)
von: Li, Ziyu, et al.
Veröffentlicht: (2024)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
AI-Assisted Requirements Engineering: An Empirical Evaluation Relative to Expert Judgment
von: Levy, Oz, et al.
Veröffentlicht: (2026)
von: Levy, Oz, et al.
Veröffentlicht: (2026)
Read, Extract, Classify: A Tool for Smarter Requirements Engineering
von: Bhattacharya, Paheli, et al.
Veröffentlicht: (2026)
von: Bhattacharya, Paheli, et al.
Veröffentlicht: (2026)
ToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems
von: Sigdel, Akshey, et al.
Veröffentlicht: (2026)
von: Sigdel, Akshey, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis
von: Ravishankara, Mayank
Veröffentlicht: (2026) -
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
von: Galimzyanov, Timur, et al.
Veröffentlicht: (2024) -
Cost-Efficient Prompt Engineering for Unsupervised Entity Resolution
von: Nananukul, Navapat, et al.
Veröffentlicht: (2023) -
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
von: Jiang, Shufan, et al.
Veröffentlicht: (2026) -
LLMs: A Game-Changer for Software Engineers?
von: Haque, Md Asraful
Veröffentlicht: (2024)