Technique to Baseline QE Artefact Generation Aligned to Quality Metrics
Fuente:
arXiv
Saved in:
| Main Authors: | Farchi, Eitan, Nayak, Kiran, Majumdar, Papia Ghosh, Route, Saritha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quality Engineering for Agile and DevOps on the Cloud and Edge
by: Farchi, Eitan, et al.
Published: (2023)
by: Farchi, Eitan, et al.
Published: (2023)
Enhancing Formal Software Specification with Artificial Intelligence
by: Nassar, Antonio Abu, et al.
Published: (2026)
by: Nassar, Antonio Abu, et al.
Published: (2026)
An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
by: Zadorojniy, Alexander, et al.
Published: (2025)
by: Zadorojniy, Alexander, et al.
Published: (2025)
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
by: Dreyfuss, Itay, et al.
Published: (2025)
by: Dreyfuss, Itay, et al.
Published: (2025)
Effective Technical Reviews
by: Ballentine, Scott, et al.
Published: (2024)
by: Ballentine, Scott, et al.
Published: (2024)
A Practical Approach to Combinatorial Test Design
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
by: Paul, Debalina Ghosh, et al.
Published: (2024)
by: Paul, Debalina Ghosh, et al.
Published: (2024)
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Structure-Aware Corpus Construction and User-Perception-Aligned Metrics for Large-Language-Model Code Completion
by: Liu, Dengfeng, et al.
Published: (2025)
by: Liu, Dengfeng, et al.
Published: (2025)
Generalized Coverage Criteria for Combinatorial Sequence Testing
by: Elyasaf, Achiya, et al.
Published: (2022)
by: Elyasaf, Achiya, et al.
Published: (2022)
GEMS: Generative Expert Metric System through Iterative Prompt Priming
by: Cheng, Ti-Chung, et al.
Published: (2024)
by: Cheng, Ti-Chung, et al.
Published: (2024)
Bridging LLM-Generated Code and Requirements: Reverse Generation technique and SBC Metric for Developer Insights
by: Ponnusamy, Ahilan Ayyachamy Nadar
Published: (2025)
by: Ponnusamy, Ahilan Ayyachamy Nadar
Published: (2025)
AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code Completion
by: Jiang, Tianyue, et al.
Published: (2026)
by: Jiang, Tianyue, et al.
Published: (2026)
Hallucinations in Code Change to Natural Language Generation: Prevalence and Evaluation of Detection Metrics
by: Liu, Chunhua, et al.
Published: (2025)
by: Liu, Chunhua, et al.
Published: (2025)
Investigating The Smells of LLM Generated Code
by: Paul, Debalina Ghosh, et al.
Published: (2025)
by: Paul, Debalina Ghosh, et al.
Published: (2025)
Enhancing LLM-Based Code Generation with Complexity Metrics: A Feedback-Driven Approach
by: Sepidband, Melika, et al.
Published: (2025)
by: Sepidband, Melika, et al.
Published: (2025)
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
by: Liu, Chenxu, et al.
Published: (2026)
by: Liu, Chenxu, et al.
Published: (2026)
Selection of Prompt Engineering Techniques for Code Generation through Predicting Code Complexity
by: Wang, Chung-Yu, et al.
Published: (2024)
by: Wang, Chung-Yu, et al.
Published: (2024)
Learning to Align Human Code Preferences
by: Yin, Xin, et al.
Published: (2025)
by: Yin, Xin, et al.
Published: (2025)
Can LLMs Generate User Stories and Assess Their Quality?
by: Quattrocchi, Giovanni, et al.
Published: (2025)
by: Quattrocchi, Giovanni, et al.
Published: (2025)
ScenEval: A Benchmark for Scenario-Based Evaluation of Code Generation
by: Paul, Debalina Ghosh, et al.
Published: (2024)
by: Paul, Debalina Ghosh, et al.
Published: (2024)
Generating Minimalist Adversarial Perturbations to Test Object-Detection Models: An Adaptive Multi-Metric Evolutionary Search Approach
by: McIntyre-Garcia, Cristopher, et al.
Published: (2024)
by: McIntyre-Garcia, Cristopher, et al.
Published: (2024)
Quality and Security Signals in AI-Generated Python Refactoring Pull Requests
by: Almukhtar, Mohamed, et al.
Published: (2026)
by: Almukhtar, Mohamed, et al.
Published: (2026)
Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code
by: He, Kaifeng, et al.
Published: (2026)
by: He, Kaifeng, et al.
Published: (2026)
Factors Influencing the Quality of AI-Generated Code: A Synthesis of Empirical Evidence
by: Geruslu, Vehid, et al.
Published: (2026)
by: Geruslu, Vehid, et al.
Published: (2026)
Structurally Aligned Subtask-Level Memory for Software Engineering Agents
by: Shen, Kangning, et al.
Published: (2026)
by: Shen, Kangning, et al.
Published: (2026)
RiskBridge: Turning CVEs into Business-Aligned Patch Priorities
by: Sheikh, Yelena Mujibur, et al.
Published: (2026)
by: Sheikh, Yelena Mujibur, et al.
Published: (2026)
Semantically Aligned Question and Code Generation for Automated Insight Generation
by: Singha, Ananya, et al.
Published: (2024)
by: Singha, Ananya, et al.
Published: (2024)
Generating High-Quality Datasets for Code Editing via Open-Source Language Models
by: Zhang, Zekai, et al.
Published: (2025)
by: Zhang, Zekai, et al.
Published: (2025)
Enhancing High-Quality Code Generation in Large Language Models with Comparative Prefix-Tuning
by: Jiang, Yuan, et al.
Published: (2025)
by: Jiang, Yuan, et al.
Published: (2025)
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
by: Kaliyev, Alibek T., et al.
Published: (2026)
by: Kaliyev, Alibek T., et al.
Published: (2026)
CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging
by: Li, Zhenyu, et al.
Published: (2026)
by: Li, Zhenyu, et al.
Published: (2026)
Predicting the Understandability of Computational Notebooks through Code Metrics Analysis
by: Ghahfarokhi, Mojtaba Mostafavi, et al.
Published: (2024)
by: Ghahfarokhi, Mojtaba Mostafavi, et al.
Published: (2024)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
by: Gandhi, Shubham, et al.
Published: (2025)
by: Gandhi, Shubham, et al.
Published: (2025)
An Empirical Study of AI Techniques in Mobile Applications
by: Li, Yinghua, et al.
Published: (2022)
by: Li, Yinghua, et al.
Published: (2022)
An Evaluation of Context Length Extrapolation in Long Code via Positional Embeddings and Efficient Attention
by: Ghosh, Madhusudan, et al.
Published: (2026)
by: Ghosh, Madhusudan, et al.
Published: (2026)
Similar Items
-
Quality Engineering for Agile and DevOps on the Cloud and Edge
by: Farchi, Eitan, et al.
Published: (2023) -
Enhancing Formal Software Specification with Artificial Intelligence
by: Nassar, Antonio Abu, et al.
Published: (2026) -
An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
by: Zadorojniy, Alexander, et al.
Published: (2025) -
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025) -
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)