PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
Fuente:
arXiv
Guardado en:
| Autor principal: | Ravishankara, Mayank |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis
por: Ravishankara, Mayank
Publicado: (2026)
por: Ravishankara, Mayank
Publicado: (2026)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
por: Galimzyanov, Timur, et al.
Publicado: (2024)
por: Galimzyanov, Timur, et al.
Publicado: (2024)
Cost-Efficient Prompt Engineering for Unsupervised Entity Resolution
por: Nananukul, Navapat, et al.
Publicado: (2023)
por: Nananukul, Navapat, et al.
Publicado: (2023)
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
por: Jiang, Shufan, et al.
Publicado: (2026)
por: Jiang, Shufan, et al.
Publicado: (2026)
LLMs: A Game-Changer for Software Engineers?
por: Haque, Md Asraful
Publicado: (2024)
por: Haque, Md Asraful
Publicado: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
por: Wang, Ruiqi, et al.
Publicado: (2025)
por: Wang, Ruiqi, et al.
Publicado: (2025)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
por: Ngassom, Sylvain Kouemo, et al.
Publicado: (2024)
por: Ngassom, Sylvain Kouemo, et al.
Publicado: (2024)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2026)
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2026)
Analysis of LLMs vs Human Experts in Requirements Engineering
por: Hymel, Cory, et al.
Publicado: (2025)
por: Hymel, Cory, et al.
Publicado: (2025)
Evaluating LLMs for Visualization Tasks
por: Khan, Saadiq Rauf, et al.
Publicado: (2025)
por: Khan, Saadiq Rauf, et al.
Publicado: (2025)
LLMs for Engineering: Teaching Models to Design High Powered Rockets
por: Simonds, Toby
Publicado: (2025)
por: Simonds, Toby
Publicado: (2025)
Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives
por: Khati, Dipin, et al.
Publicado: (2025)
por: Khati, Dipin, et al.
Publicado: (2025)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
por: Wu, Fan, et al.
Publicado: (2026)
por: Wu, Fan, et al.
Publicado: (2026)
Lifecycle-Aware code generation: Leveraging Software Engineering Phases in LLMs
por: Xing, Xing, et al.
Publicado: (2025)
por: Xing, Xing, et al.
Publicado: (2025)
Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research
por: Trinkenreich, Bianca, et al.
Publicado: (2025)
por: Trinkenreich, Bianca, et al.
Publicado: (2025)
Evaluating the Energy-Efficiency of the Code Generated by LLMs
por: Islam, Md Arman, et al.
Publicado: (2025)
por: Islam, Md Arman, et al.
Publicado: (2025)
Evaluating the Generalizability of LLMs in Automated Program Repair
por: Li, Fengjie, et al.
Publicado: (2025)
por: Li, Fengjie, et al.
Publicado: (2025)
From Inductive to Deductive: LLMs-Based Qualitative Data Analysis in Requirements Engineering
por: Shah, Syed Tauhid Ullah, et al.
Publicado: (2025)
por: Shah, Syed Tauhid Ullah, et al.
Publicado: (2025)
Holistic Evaluation of State-of-the-Art LLMs for Code Generation
por: Zhang, Le, et al.
Publicado: (2025)
por: Zhang, Le, et al.
Publicado: (2025)
Using LLMs in Software Requirements Specifications: An Empirical Evaluation
por: Krishna, Madhava, et al.
Publicado: (2024)
por: Krishna, Madhava, et al.
Publicado: (2024)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
por: Wang, Hanbin, et al.
Publicado: (2025)
por: Wang, Hanbin, et al.
Publicado: (2025)
Repository Intelligence Graph: Deterministic Architectural Map for LLM Code Assistants
por: Cherny-Shahar, Tsvi, et al.
Publicado: (2026)
por: Cherny-Shahar, Tsvi, et al.
Publicado: (2026)
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
por: Pan, Chenkai, et al.
Publicado: (2026)
por: Pan, Chenkai, et al.
Publicado: (2026)
TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance
por: Bruches, Elena, et al.
Publicado: (2026)
por: Bruches, Elena, et al.
Publicado: (2026)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
por: Singha, Ananya, et al.
Publicado: (2025)
por: Singha, Ananya, et al.
Publicado: (2025)
FullStack Bench: Evaluating LLMs as Full Stack Coders
por: Bytedance-Seed-Foundation-Code-Team, et al.
Publicado: (2024)
por: Bytedance-Seed-Foundation-Code-Team, et al.
Publicado: (2024)
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
por: Li, Jingyue, et al.
Publicado: (2026)
por: Li, Jingyue, et al.
Publicado: (2026)
ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation
por: Xianpeng, et al.
Publicado: (2026)
por: Xianpeng, et al.
Publicado: (2026)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
por: Zhu, Hongda, et al.
Publicado: (2025)
por: Zhu, Hongda, et al.
Publicado: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
por: Fandina, Ora Nova, et al.
Publicado: (2025)
por: Fandina, Ora Nova, et al.
Publicado: (2025)
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
por: Khati, Dipin, et al.
Publicado: (2026)
por: Khati, Dipin, et al.
Publicado: (2026)
An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains
por: Yan, Zihe, et al.
Publicado: (2025)
por: Yan, Zihe, et al.
Publicado: (2025)
Adaptive Hierarchical Evaluation of LLMs and SAST tools for CWE Prediction in Python
por: Adnan, Muntasir, et al.
Publicado: (2026)
por: Adnan, Muntasir, et al.
Publicado: (2026)
Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects
por: Nunes, Henrique, et al.
Publicado: (2025)
por: Nunes, Henrique, et al.
Publicado: (2025)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
por: Wang, Shufan, et al.
Publicado: (2025)
por: Wang, Shufan, et al.
Publicado: (2025)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
por: Li, Ziyu, et al.
Publicado: (2024)
por: Li, Ziyu, et al.
Publicado: (2024)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
por: Gao, Shuzheng, et al.
Publicado: (2025)
por: Gao, Shuzheng, et al.
Publicado: (2025)
AI-Assisted Requirements Engineering: An Empirical Evaluation Relative to Expert Judgment
por: Levy, Oz, et al.
Publicado: (2026)
por: Levy, Oz, et al.
Publicado: (2026)
Read, Extract, Classify: A Tool for Smarter Requirements Engineering
por: Bhattacharya, Paheli, et al.
Publicado: (2026)
por: Bhattacharya, Paheli, et al.
Publicado: (2026)
ToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems
por: Sigdel, Akshey, et al.
Publicado: (2026)
por: Sigdel, Akshey, et al.
Publicado: (2026)
Ejemplares similares
-
CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis
por: Ravishankara, Mayank
Publicado: (2026) -
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
por: Galimzyanov, Timur, et al.
Publicado: (2024) -
Cost-Efficient Prompt Engineering for Unsupervised Entity Resolution
por: Nananukul, Navapat, et al.
Publicado: (2023) -
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
por: Jiang, Shufan, et al.
Publicado: (2026) -
LLMs: A Game-Changer for Software Engineers?
por: Haque, Md Asraful
Publicado: (2024)