Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
Fuente:
arXiv
Saved in:
| Main Authors: | Fandina, Ora Nova, Farchi, Eitan, Froimovich, Shmulik, Katan, Rami, Podolsky, Alice, Raz, Orna, Ziv, Avi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025)
by: Froimovich, Shmulik, et al.
Published: (2025)
LaajMeter: A Framework for LaaJ Evaluation
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
by: Dreyfuss, Itay, et al.
Published: (2025)
by: Dreyfuss, Itay, et al.
Published: (2025)
Uncovering Code Insights: Leveraging GitHub Artifacts for Deeper Code Understanding
by: Nevo, Ziv, et al.
Published: (2025)
by: Nevo, Ziv, et al.
Published: (2025)
Using Combinatorial Optimization to Design a High quality LLM Solution
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
Quality Engineering for Agile and DevOps on the Cloud and Edge
by: Farchi, Eitan, et al.
Published: (2023)
by: Farchi, Eitan, et al.
Published: (2023)
Enhancing Formal Software Specification with Artificial Intelligence
by: Nassar, Antonio Abu, et al.
Published: (2026)
by: Nassar, Antonio Abu, et al.
Published: (2026)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
by: Fandina, Ora Nova, et al.
Published: (2024)
by: Fandina, Ora Nova, et al.
Published: (2024)
Effective Technical Reviews
by: Ballentine, Scott, et al.
Published: (2024)
by: Ballentine, Scott, et al.
Published: (2024)
A Practical Approach to Combinatorial Test Design
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
by: Zadorojniy, Alexander, et al.
Published: (2025)
by: Zadorojniy, Alexander, et al.
Published: (2025)
Black-Box Bug-Amplification for Multithreaded Software
by: Weiss, Yeshayahu, et al.
Published: (2025)
by: Weiss, Yeshayahu, et al.
Published: (2025)
Evaluating perturbation robustness of generative systems that use COBOL code inputs
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Rethinking Artifact Evaluation for Software Engineering in the Age of Generative AI
by: Treude, Christoph, et al.
Published: (2026)
by: Treude, Christoph, et al.
Published: (2026)
Automated Personnel Selection for Software Engineers Using LLM-Based Profile Evaluation
by: Karim, Ahmed Akib Jawad, et al.
Published: (2024)
by: Karim, Ahmed Akib Jawad, et al.
Published: (2024)
Generalized Coverage Criteria for Combinatorial Sequence Testing
by: Elyasaf, Achiya, et al.
Published: (2022)
by: Elyasaf, Achiya, et al.
Published: (2022)
Agent-Based Software Artifact Evaluation
by: Wu, Zhaonan, et al.
Published: (2026)
by: Wu, Zhaonan, et al.
Published: (2026)
Research Artifacts in Software Engineering Publications: Status and Trends
by: Liu, Mugeng, et al.
Published: (2024)
by: Liu, Mugeng, et al.
Published: (2024)
From Issues to Insights: RAG-based Explanation Generation from Software Engineering Artifacts
by: Pöttgen, Daniel, et al.
Published: (2026)
by: Pöttgen, Daniel, et al.
Published: (2026)
Technique to Baseline QE Artefact Generation Aligned to Quality Metrics
by: Farchi, Eitan, et al.
Published: (2025)
by: Farchi, Eitan, et al.
Published: (2025)
Evaluating LLM Agents on Automated Software Analysis Tasks
by: Bouzenia, Islem, et al.
Published: (2026)
by: Bouzenia, Islem, et al.
Published: (2026)
Prompts as Software Engineering Artifacts: A Research Agenda and Preliminary Findings
by: Villamizar, Hugo, et al.
Published: (2025)
by: Villamizar, Hugo, et al.
Published: (2025)
Research Artifacts in Secondary Studies: A Systematic Mapping in Software Engineering
by: Huotala, Aleksi, et al.
Published: (2025)
by: Huotala, Aleksi, et al.
Published: (2025)
BASFuzz: Towards Robustness Evaluation of LLM-based NLP Software via Automated Fuzz Testing
by: Xiao, Mingxuan, et al.
Published: (2025)
by: Xiao, Mingxuan, et al.
Published: (2025)
Is Your Automated Software Engineer Trustworthy?
by: Mathews, Noble Saji, et al.
Published: (2025)
by: Mathews, Noble Saji, et al.
Published: (2025)
Automated Quantum Software and AI Engineering
by: Siavash, Nazanin, et al.
Published: (2026)
by: Siavash, Nazanin, et al.
Published: (2026)
The State of Open Science in Software Engineering Research: A Case Study of ICSE Artifacts
by: Muttakin, Al, et al.
Published: (2026)
by: Muttakin, Al, et al.
Published: (2026)
Operationalizing Software Engineering Theories for Practical Validation
by: Alves, Isaque, et al.
Published: (2026)
by: Alves, Isaque, et al.
Published: (2026)
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
by: Kaliyev, Alibek T., et al.
Published: (2026)
by: Kaliyev, Alibek T., et al.
Published: (2026)
On Developing an Artifact-based Approach to Regulatory Requirements Engineering
by: Kosenkov, Oleksandr, et al.
Published: (2024)
by: Kosenkov, Oleksandr, et al.
Published: (2024)
Assessing the Robustness of LLM-based NLP Software via Automated Testing
by: Xiao, Mingxuan, et al.
Published: (2024)
by: Xiao, Mingxuan, et al.
Published: (2024)
How do Software Engineering Researchers Use GitHub? An Empirical Study of Artifacts & Impact
by: Alrashedy, Kamel, et al.
Published: (2023)
by: Alrashedy, Kamel, et al.
Published: (2023)
Reporting LLM Prompting in Automated Software Engineering: A Guideline Based on Current Practices and Expectations
by: Korn, Alexander, et al.
Published: (2026)
by: Korn, Alexander, et al.
Published: (2026)
Integrating Various Software Artifacts for Better LLM-based Bug Localization and Program Repair
by: Feng, Qiong, et al.
Published: (2024)
by: Feng, Qiong, et al.
Published: (2024)
Assured LLM-Based Software Engineering
by: Alshahwan, Nadia, et al.
Published: (2024)
by: Alshahwan, Nadia, et al.
Published: (2024)
Measuring LLM Trust Allocation Across Conflicting Software Artifacts
by: Ulfat, Noshin, et al.
Published: (2026)
by: Ulfat, Noshin, et al.
Published: (2026)
Evaluation of LLM-Based Software Engineering Tools: Practices, Challenges, and Future Directions
by: Torun, Utku Boran, et al.
Published: (2026)
by: Torun, Utku Boran, et al.
Published: (2026)
Similar Items
-
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024) -
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025) -
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
by: Fandina, Ora Nova, et al.
Published: (2025) -
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025) -
LaajMeter: A Framework for LaaJ Evaluation
by: Ackerman, Samuel, et al.
Published: (2025)