Evaluating the effectiveness of LLM-based interoperability
Fuente:
arXiv
Saved in:
| Main Authors: | Falcão, Rodrigo, Schweitzer, Stefan, Siebert, Julien, Calvet, Emily, Elberzhager, Frank |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward architecting self-coding information systems
by: Falcão, Rodrigo, et al.
Published: (2026)
by: Falcão, Rodrigo, et al.
Published: (2026)
Experimentation package for evaluating the effectiveness of LLM-based interoperability
by: Falcão, Rodrigo, et al.
Published: (2025)
by: Falcão, Rodrigo, et al.
Published: (2025)
Experiences in Using the V-Model as a Framework for Applied Doctoral Research
by: Falcão, Rodrigo, et al.
Published: (2024)
by: Falcão, Rodrigo, et al.
Published: (2024)
Using LLMs to Evaluate Architecture Documents: Results from a Digital Marketplace Environment
by: Elberzhager, Frank, et al.
Published: (2026)
by: Elberzhager, Frank, et al.
Published: (2026)
Causal Software Engineering: A Vision and Roadmap
by: Pietrantuono, Roberto, et al.
Published: (2026)
by: Pietrantuono, Roberto, et al.
Published: (2026)
Evaluation of large language models for assessing code maintainability
by: Dillmann, Marc, et al.
Published: (2024)
by: Dillmann, Marc, et al.
Published: (2024)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Engineering a sustainable world by enhancing the scope of systems of systems engineering and mastering dynamics
by: Adler, Rasmus, et al.
Published: (2024)
by: Adler, Rasmus, et al.
Published: (2024)
RESTestBench: A Benchmark for Evaluating the Effectiveness of LLM-Generated REST API Test Cases from NL Requirements
by: Kogler, Leon, et al.
Published: (2026)
by: Kogler, Leon, et al.
Published: (2026)
An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains
by: Yan, Zihe, et al.
Published: (2025)
by: Yan, Zihe, et al.
Published: (2025)
Rubric Is All You Need: Enhancing LLM-based Code Evaluation With Question-Specific Rubrics
by: Pathak, Aditya, et al.
Published: (2025)
by: Pathak, Aditya, et al.
Published: (2025)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
by: Agarwal, Anisha, et al.
Published: (2024)
by: Agarwal, Anisha, et al.
Published: (2024)
From research to clinic: Accelerating the translation of clinical decision support systems by making synthetic data interoperable
by: Chauhan, Pavitra, et al.
Published: (2023)
by: Chauhan, Pavitra, et al.
Published: (2023)
Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution
by: Yu, Kai, et al.
Published: (2026)
by: Yu, Kai, et al.
Published: (2026)
LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile Apps
by: Zhao, Shanhui, et al.
Published: (2025)
by: Zhao, Shanhui, et al.
Published: (2025)
Multimodal Approach for Harmonized System Code Prediction
by: Amel, Otmane, et al.
Published: (2024)
by: Amel, Otmane, et al.
Published: (2024)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
by: Pan, Zhiyuan, et al.
Published: (2025)
by: Pan, Zhiyuan, et al.
Published: (2025)
Evaluating LLM-Based Test Generation Under Software Evolution
by: Haroon, Sabaat, et al.
Published: (2026)
by: Haroon, Sabaat, et al.
Published: (2026)
LLM-based Iterative Approach to Metamodeling in Automotive
by: Petrovic, Nenad, et al.
Published: (2025)
by: Petrovic, Nenad, et al.
Published: (2025)
Uncertainty Quantification for LLM-based Code Generation
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
Enabling Predictive Maintenance in District Heating Substations: A Labelled Dataset and Fault Detection Evaluation Framework based on Service Data
by: Roelofs, Cyriana M. A., et al.
Published: (2025)
by: Roelofs, Cyriana M. A., et al.
Published: (2025)
Evolving Excellence: Automated Optimization of LLM-based Agents
by: Brookes, Paul, et al.
Published: (2025)
by: Brookes, Paul, et al.
Published: (2025)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
SWE-QA: A Dataset and Benchmark for Complex Code Understanding
by: Elkoussy, Laïla, et al.
Published: (2026)
by: Elkoussy, Laïla, et al.
Published: (2026)
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
by: Yan, Shuo, et al.
Published: (2025)
by: Yan, Shuo, et al.
Published: (2025)
LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
Tricky$^2$: Towards a Benchmark for Evaluating Human and LLM Error Interactions
by: Granger, Cole, et al.
Published: (2026)
by: Granger, Cole, et al.
Published: (2026)
PyBench: Evaluating LLM Agent on various real-world coding tasks
by: Zhang, Yaolun, et al.
Published: (2024)
by: Zhang, Yaolun, et al.
Published: (2024)
AISysRev -- LLM-based Tool for Title-abstract Screening
by: Huotala, Aleksi, et al.
Published: (2025)
by: Huotala, Aleksi, et al.
Published: (2025)
An Empirical Study on LLM-based Agents for Automated Bug Fixing
by: Meng, Xiangxin, et al.
Published: (2024)
by: Meng, Xiangxin, et al.
Published: (2024)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
by: Huang, Donghao, et al.
Published: (2025)
by: Huang, Donghao, et al.
Published: (2025)
Advancing Automated Ethical Profiling in SE: a Zero-Shot Evaluation of LLM Reasoning
by: Migliarini, Patrizio, et al.
Published: (2025)
by: Migliarini, Patrizio, et al.
Published: (2025)
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
by: Li, Yuanyang, et al.
Published: (2026)
by: Li, Yuanyang, et al.
Published: (2026)
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
by: Kaliyev, Alibek T., et al.
Published: (2026)
by: Kaliyev, Alibek T., et al.
Published: (2026)
MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems
by: Jia, Jin, et al.
Published: (2026)
by: Jia, Jin, et al.
Published: (2026)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
by: Tóth, Rebeka, et al.
Published: (2024)
by: Tóth, Rebeka, et al.
Published: (2024)
Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
by: Xia, Boming, et al.
Published: (2024)
by: Xia, Boming, et al.
Published: (2024)
Evaluating Agent-based Program Repair at Google
by: Rondon, Pat, et al.
Published: (2025)
by: Rondon, Pat, et al.
Published: (2025)
Similar Items
-
Toward architecting self-coding information systems
by: Falcão, Rodrigo, et al.
Published: (2026) -
Experimentation package for evaluating the effectiveness of LLM-based interoperability
by: Falcão, Rodrigo, et al.
Published: (2025) -
Experiences in Using the V-Model as a Framework for Applied Doctoral Research
by: Falcão, Rodrigo, et al.
Published: (2024) -
Using LLMs to Evaluate Architecture Documents: Results from a Digital Marketplace Environment
by: Elberzhager, Frank, et al.
Published: (2026) -
Causal Software Engineering: A Vision and Roadmap
by: Pietrantuono, Roberto, et al.
Published: (2026)