GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Shufan, Chen, Chios, Chen, Zhiyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
von: Wang, Shufan, et al.
Veröffentlicht: (2025)
von: Wang, Shufan, et al.
Veröffentlicht: (2025)
LLMs: A Game-Changer for Software Engineers?
von: Haque, Md Asraful
Veröffentlicht: (2024)
von: Haque, Md Asraful
Veröffentlicht: (2024)
On Unified Prompt Tuning for Request Quality Assurance in Public Code Review
von: Chen, Xinyu, et al.
Veröffentlicht: (2024)
von: Chen, Xinyu, et al.
Veröffentlicht: (2024)
Knowledge-Guided Prompt Learning for Request Quality Assurance in Public Code Review
von: Li, Lin, et al.
Veröffentlicht: (2024)
von: Li, Lin, et al.
Veröffentlicht: (2024)
Causal Reasoning in Software Quality Assurance: A Systematic Review
von: Giamattei, Luca, et al.
Veröffentlicht: (2024)
von: Giamattei, Luca, et al.
Veröffentlicht: (2024)
Quality Assurance of LLM-generated Code: Addressing Non-Functional Quality Characteristics
von: Sun, Xin, et al.
Veröffentlicht: (2025)
von: Sun, Xin, et al.
Veröffentlicht: (2025)
CoDefeater: Using LLMs To Find Defeaters in Assurance Cases
von: Gohar, Usman, et al.
Veröffentlicht: (2024)
von: Gohar, Usman, et al.
Veröffentlicht: (2024)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
von: Singha, Ananya, et al.
Veröffentlicht: (2025)
von: Singha, Ananya, et al.
Veröffentlicht: (2025)
ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation
von: Xianpeng, et al.
Veröffentlicht: (2026)
von: Xianpeng, et al.
Veröffentlicht: (2026)
WALL: A Web Application for Automated Quality Assurance using Large Language Models
von: Abtahi, Seyed Moein, et al.
Veröffentlicht: (2025)
von: Abtahi, Seyed Moein, et al.
Veröffentlicht: (2025)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2026)
von: Rodriguez-Cardenas, Daniel, et al.
Veröffentlicht: (2026)
The Rise of Agentic Testing: Multi-Agent Systems for Robust Software Quality Assurance
von: Naqvi, Saba, et al.
Veröffentlicht: (2026)
von: Naqvi, Saba, et al.
Veröffentlicht: (2026)
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
von: Koohestani, Roham, et al.
Veröffentlicht: (2025)
von: Koohestani, Roham, et al.
Veröffentlicht: (2025)
AI-Driven Tools in Modern Software Quality Assurance: An Assessment of Benefits, Challenges, and Future Directions
von: Pysmennyi, Ihor, et al.
Veröffentlicht: (2025)
von: Pysmennyi, Ihor, et al.
Veröffentlicht: (2025)
Lessons Learned from the Use of Generative AI in Engineering and Quality Assurance of a WEB System for Healthcare
von: Travassos, Guilherme H., et al.
Veröffentlicht: (2025)
von: Travassos, Guilherme H., et al.
Veröffentlicht: (2025)
Argus: Resilience-Oriented Safety Assurance Framework for End-to-End ADSs
von: Wang, Dingji, et al.
Veröffentlicht: (2025)
von: Wang, Dingji, et al.
Veröffentlicht: (2025)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
von: Zhu, Hongda, et al.
Veröffentlicht: (2025)
von: Zhu, Hongda, et al.
Veröffentlicht: (2025)
Position: Early-Stage Quality Assurance in Annotation Pipelines Is More Cost-Effective Than Late-Stage Validation
von: Kothari, Sunil, et al.
Veröffentlicht: (2026)
von: Kothari, Sunil, et al.
Veröffentlicht: (2026)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs
von: Momcilovic, Tomas Bueno, et al.
Veröffentlicht: (2024)
von: Momcilovic, Tomas Bueno, et al.
Veröffentlicht: (2024)
PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
von: Ravishankara, Mayank
Veröffentlicht: (2026)
von: Ravishankara, Mayank
Veröffentlicht: (2026)
Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs
von: Chen, Zhiyang, et al.
Veröffentlicht: (2025)
von: Chen, Zhiyang, et al.
Veröffentlicht: (2025)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
von: Wu, Fan, et al.
Veröffentlicht: (2026)
von: Wu, Fan, et al.
Veröffentlicht: (2026)
Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code
von: He, Kaifeng, et al.
Veröffentlicht: (2026)
von: He, Kaifeng, et al.
Veröffentlicht: (2026)
Evaluating the Generalizability of LLMs in Automated Program Repair
von: Li, Fengjie, et al.
Veröffentlicht: (2025)
von: Li, Fengjie, et al.
Veröffentlicht: (2025)
AI Assurance: A Comprehensive Testing Strategy for Enterprise AI Systems
von: Badagi, Chitra, et al.
Veröffentlicht: (2026)
von: Badagi, Chitra, et al.
Veröffentlicht: (2026)
Skeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation
von: Zhang, Xing, et al.
Veröffentlicht: (2025)
von: Zhang, Xing, et al.
Veröffentlicht: (2025)
CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval
von: Geng, Jiahui, et al.
Veröffentlicht: (2025)
von: Geng, Jiahui, et al.
Veröffentlicht: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
CodeRepoQA: A Large-scale Benchmark for Software Engineering Question Answering
von: Hu, Ruida, et al.
Veröffentlicht: (2024)
von: Hu, Ruida, et al.
Veröffentlicht: (2024)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
FullStack Bench: Evaluating LLMs as Full Stack Coders
von: Bytedance-Seed-Foundation-Code-Team, et al.
Veröffentlicht: (2024)
von: Bytedance-Seed-Foundation-Code-Team, et al.
Veröffentlicht: (2024)
Automated Population-Level Audit Assurance via AI-Based Document Intelligence
von: Vasudevan, Santosh, et al.
Veröffentlicht: (2026)
von: Vasudevan, Santosh, et al.
Veröffentlicht: (2026)
OntoGSN: An Ontology-Based Framework for Semantic Management and Extension of Assurance Cases
von: Momcilovic, Tomas Bueno, et al.
Veröffentlicht: (2025)
von: Momcilovic, Tomas Bueno, et al.
Veröffentlicht: (2025)
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
von: Pan, Chenkai, et al.
Veröffentlicht: (2026)
von: Pan, Chenkai, et al.
Veröffentlicht: (2026)
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
von: Lian, Keke, et al.
Veröffentlicht: (2025)
von: Lian, Keke, et al.
Veröffentlicht: (2025)
A Note on Code Quality Score: LLMs for Maintainable Large Codebases
von: Wong, Sherman, et al.
Veröffentlicht: (2025)
von: Wong, Sherman, et al.
Veröffentlicht: (2025)
Can LLMs Generate User Stories and Assess Their Quality?
von: Quattrocchi, Giovanni, et al.
Veröffentlicht: (2025)
von: Quattrocchi, Giovanni, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
von: Wang, Shufan, et al.
Veröffentlicht: (2025) -
LLMs: A Game-Changer for Software Engineers?
von: Haque, Md Asraful
Veröffentlicht: (2024) -
On Unified Prompt Tuning for Request Quality Assurance in Public Code Review
von: Chen, Xinyu, et al.
Veröffentlicht: (2024) -
Knowledge-Guided Prompt Learning for Request Quality Assurance in Public Code Review
von: Li, Lin, et al.
Veröffentlicht: (2024) -
Causal Reasoning in Software Quality Assurance: A Systematic Review
von: Giamattei, Luca, et al.
Veröffentlicht: (2024)