Methodological Framework for Quantifying Semantic Test Coverage in RAG Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Broestl, Noah, Abdalla, Adel Nasser, Bale, Rajprakash, Gupta, Hersh, Struever, Max |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification
di: Xu, Jiacheng, et al.
Pubblicazione: (2025)
di: Xu, Jiacheng, et al.
Pubblicazione: (2025)
Exploring the Integration of Large Language Models in Industrial Test Maintenance Processes
di: Liu, Jingxiong, et al.
Pubblicazione: (2024)
di: Liu, Jingxiong, et al.
Pubblicazione: (2024)
Agile Story-Point Estimation: Is RAG a Better Way to Go?
di: Maha, Lamyea, et al.
Pubblicazione: (2026)
di: Maha, Lamyea, et al.
Pubblicazione: (2026)
Uncovering Discrimination Clusters: Quantifying and Explaining Systematic Fairness Violations
di: Akash, Ranit Debnath, et al.
Pubblicazione: (2025)
di: Akash, Ranit Debnath, et al.
Pubblicazione: (2025)
A Conceptual Framework for Ethical Evaluation of Machine Learning Systems
di: Gupta, Neha R., et al.
Pubblicazione: (2024)
di: Gupta, Neha R., et al.
Pubblicazione: (2024)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
di: Han, Xiaoke, et al.
Pubblicazione: (2025)
di: Han, Xiaoke, et al.
Pubblicazione: (2025)
Codehacks: A Dataset of Adversarial Tests for Competitive Programming Problems Obtained from Codeforces
di: Hort, Max, et al.
Pubblicazione: (2025)
di: Hort, Max, et al.
Pubblicazione: (2025)
Semantic-Preserving Transformations as Mutation Operators: A Study on Their Effectiveness in Defect Detection
di: Hort, Max, et al.
Pubblicazione: (2025)
di: Hort, Max, et al.
Pubblicazione: (2025)
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
di: Samsonau, Sergey V.
Pubblicazione: (2026)
di: Samsonau, Sergey V.
Pubblicazione: (2026)
Can Search-Based Testing with Pareto Optimization Effectively Cover Failure-Revealing Test Inputs?
di: Sorokin, Lev, et al.
Pubblicazione: (2024)
di: Sorokin, Lev, et al.
Pubblicazione: (2024)
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
di: Wei, Zeming, et al.
Pubblicazione: (2026)
di: Wei, Zeming, et al.
Pubblicazione: (2026)
How Robustly do LLMs Understand Execution Semantics?
di: Spiess, Claudio, et al.
Pubblicazione: (2026)
di: Spiess, Claudio, et al.
Pubblicazione: (2026)
When Your LLM Reaches End-of-Life: A Framework for Confident Model Migration in Production Systems
di: Casey, Emma, et al.
Pubblicazione: (2026)
di: Casey, Emma, et al.
Pubblicazione: (2026)
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
di: Fatima, Sakina, et al.
Pubblicazione: (2023)
di: Fatima, Sakina, et al.
Pubblicazione: (2023)
Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation
di: Jacopin, Éric
Pubblicazione: (2026)
di: Jacopin, Éric
Pubblicazione: (2026)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
di: Jiang, Shan, et al.
Pubblicazione: (2026)
di: Jiang, Shan, et al.
Pubblicazione: (2026)
Generative AI to Generate Test Data Generators
di: Baudry, Benoit, et al.
Pubblicazione: (2024)
di: Baudry, Benoit, et al.
Pubblicazione: (2024)
Understanding LLM-Driven Test Oracle Generation
di: Bodicoat, Adam, et al.
Pubblicazione: (2026)
di: Bodicoat, Adam, et al.
Pubblicazione: (2026)
Code Generation by Differential Test Time Scaling
di: He, Yifeng, et al.
Pubblicazione: (2026)
di: He, Yifeng, et al.
Pubblicazione: (2026)
A Stochastic Differential Equation Framework for Multi-Objective LLM Interactions: Dynamical Systems Analysis with Code Generation Applications
di: Shukla, Shivani, et al.
Pubblicazione: (2025)
di: Shukla, Shivani, et al.
Pubblicazione: (2025)
Mutation-Guided LLM-based Test Generation at Meta
di: Foster, Christopher, et al.
Pubblicazione: (2025)
di: Foster, Christopher, et al.
Pubblicazione: (2025)
DeepKnowledge: Generalisation-Driven Deep Learning Testing
di: Missaoui, Sondess, et al.
Pubblicazione: (2024)
di: Missaoui, Sondess, et al.
Pubblicazione: (2024)
A Theoretical Analysis of Test-Driven Code Generation
di: Menet, Nicolas, et al.
Pubblicazione: (2026)
di: Menet, Nicolas, et al.
Pubblicazione: (2026)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
di: Xu, WeiZhe, et al.
Pubblicazione: (2026)
di: Xu, WeiZhe, et al.
Pubblicazione: (2026)
The Impact of Software Testing with Quantum Optimization Meets Machine Learning
di: Bandarupalli, Gopichand
Pubblicazione: (2025)
di: Bandarupalli, Gopichand
Pubblicazione: (2025)
RBT4DNN: Requirements-based Testing of Neural Networks
di: Mozumder, Nusrat Jahan, et al.
Pubblicazione: (2025)
di: Mozumder, Nusrat Jahan, et al.
Pubblicazione: (2025)
Reinforcement Learning for Online Testing of Autonomous Driving Systems: a Replication and Extension Study
di: Giamattei, Luca, et al.
Pubblicazione: (2024)
di: Giamattei, Luca, et al.
Pubblicazione: (2024)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
di: Zhou, Zenghui, et al.
Pubblicazione: (2026)
di: Zhou, Zenghui, et al.
Pubblicazione: (2026)
Using Quality Attribute Scenarios for ML Model Test Case Generation
di: Brower-Sinning, Rachel, et al.
Pubblicazione: (2024)
di: Brower-Sinning, Rachel, et al.
Pubblicazione: (2024)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
di: Bai, Yifan, et al.
Pubblicazione: (2026)
di: Bai, Yifan, et al.
Pubblicazione: (2026)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
di: Ding, Yifeng, et al.
Pubblicazione: (2026)
di: Ding, Yifeng, et al.
Pubblicazione: (2026)
Read, Extract, Classify: A Tool for Smarter Requirements Engineering
di: Bhattacharya, Paheli, et al.
Pubblicazione: (2026)
di: Bhattacharya, Paheli, et al.
Pubblicazione: (2026)
Zero-Shot Attribution for Large Language Models: A Distribution Testing Approach
di: Canonne, Clément L., et al.
Pubblicazione: (2025)
di: Canonne, Clément L., et al.
Pubblicazione: (2025)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
di: Mündler, Niels, et al.
Pubblicazione: (2024)
di: Mündler, Niels, et al.
Pubblicazione: (2024)
MIST-RL: Mutation-based Incremental Suite Testing via Reinforcement Learning
di: Zhu, Sicheng, et al.
Pubblicazione: (2026)
di: Zhu, Sicheng, et al.
Pubblicazione: (2026)
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study
di: Storhaug, André, et al.
Pubblicazione: (2024)
di: Storhaug, André, et al.
Pubblicazione: (2024)
CDS4RAG: Cyclic Dual-Sequential Hyperparameter Optimization for RAG
di: Chen, Pengzhou, et al.
Pubblicazione: (2026)
di: Chen, Pengzhou, et al.
Pubblicazione: (2026)
A Reference Architecture of Reinforcement Learning Frameworks
di: Liu, Xiaoran, et al.
Pubblicazione: (2026)
di: Liu, Xiaoran, et al.
Pubblicazione: (2026)
A Framework to Model ML Engineering Processes
di: Morales, Sergio, et al.
Pubblicazione: (2024)
di: Morales, Sergio, et al.
Pubblicazione: (2024)
AI-driven Java Performance Testing: Balancing Result Quality with Testing Time
di: Traini, Luca, et al.
Pubblicazione: (2024)
di: Traini, Luca, et al.
Pubblicazione: (2024)
Documenti analoghi
-
CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification
di: Xu, Jiacheng, et al.
Pubblicazione: (2025) -
Exploring the Integration of Large Language Models in Industrial Test Maintenance Processes
di: Liu, Jingxiong, et al.
Pubblicazione: (2024) -
Agile Story-Point Estimation: Is RAG a Better Way to Go?
di: Maha, Lamyea, et al.
Pubblicazione: (2026) -
Uncovering Discrimination Clusters: Quantifying and Explaining Systematic Fairness Violations
di: Akash, Ranit Debnath, et al.
Pubblicazione: (2025) -
A Conceptual Framework for Ethical Evaluation of Machine Learning Systems
di: Gupta, Neha R., et al.
Pubblicazione: (2024)