The Explabox: Model-Agnostic Machine Learning Transparency & Analysis
Fuente:
arXiv
Salvato in:
| Autori principali: | Robeer, Marcel, Bron, Michiel, Herrewijnen, Elize, Hoeseni, Riwish, Bex, Floris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
di: Li, Meiziniu, et al.
Pubblicazione: (2022)
di: Li, Meiziniu, et al.
Pubblicazione: (2022)
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
di: Li, Meiziniu, et al.
Pubblicazione: (2024)
di: Li, Meiziniu, et al.
Pubblicazione: (2024)
Demystifying the Silence of Correctness Bugs in PyTorch Compiler
di: Li, Meiziniu, et al.
Pubblicazione: (2026)
di: Li, Meiziniu, et al.
Pubblicazione: (2026)
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
di: More, Riddhi, et al.
Pubblicazione: (2025)
di: More, Riddhi, et al.
Pubblicazione: (2025)
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
di: More, Riddhi, et al.
Pubblicazione: (2025)
di: More, Riddhi, et al.
Pubblicazione: (2025)
Towards Explainable Test Case Prioritisation with Learning-to-Rank Models
di: Ramírez, Aurora, et al.
Pubblicazione: (2024)
di: Ramírez, Aurora, et al.
Pubblicazione: (2024)
On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations
di: Hundal, Rajdeep Singh, et al.
Pubblicazione: (2025)
di: Hundal, Rajdeep Singh, et al.
Pubblicazione: (2025)
Review Beats Planning: Dual-Model Interaction Patterns for Code Synthesis
di: Miller, Jan
Pubblicazione: (2026)
di: Miller, Jan
Pubblicazione: (2026)
Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair
di: Cambronero, José, et al.
Pubblicazione: (2025)
di: Cambronero, José, et al.
Pubblicazione: (2025)
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
di: Ahmed, Syed Yusuf, et al.
Pubblicazione: (2026)
di: Ahmed, Syed Yusuf, et al.
Pubblicazione: (2026)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
di: Chang, Hung-Fu, et al.
Pubblicazione: (2025)
di: Chang, Hung-Fu, et al.
Pubblicazione: (2025)
MeDeT: Medical Device Digital Twins Creation with Few-shot Meta-learning
di: Sartaj, Hassan, et al.
Pubblicazione: (2024)
di: Sartaj, Hassan, et al.
Pubblicazione: (2024)
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
di: Wu, Jiaqing, et al.
Pubblicazione: (2026)
di: Wu, Jiaqing, et al.
Pubblicazione: (2026)
AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
di: Bhardwaj, Varun Pratap
Pubblicazione: (2026)
di: Bhardwaj, Varun Pratap
Pubblicazione: (2026)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
di: Terragni, Valerio
Pubblicazione: (2026)
di: Terragni, Valerio
Pubblicazione: (2026)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
di: Bradbury, Jeremy S., et al.
Pubblicazione: (2024)
di: Bradbury, Jeremy S., et al.
Pubblicazione: (2024)
Experience with GitHub Copilot for Developer Productivity at Zoominfo
di: Bakal, Gal, et al.
Pubblicazione: (2025)
di: Bakal, Gal, et al.
Pubblicazione: (2025)
Understanding and Detecting Flaky Builds in GitHub Actions
di: Ge, Wenhao, et al.
Pubblicazione: (2026)
di: Ge, Wenhao, et al.
Pubblicazione: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
di: Hu, Yuelin, et al.
Pubblicazione: (2026)
di: Hu, Yuelin, et al.
Pubblicazione: (2026)
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
di: Majdinasab, Vahid, et al.
Pubblicazione: (2024)
di: Majdinasab, Vahid, et al.
Pubblicazione: (2024)
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
di: Jana, Prithwish, et al.
Pubblicazione: (2023)
di: Jana, Prithwish, et al.
Pubblicazione: (2023)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
di: Liu, Shunyu, et al.
Pubblicazione: (2025)
di: Liu, Shunyu, et al.
Pubblicazione: (2025)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
di: Rehan, Tzafrir
Pubblicazione: (2026)
di: Rehan, Tzafrir
Pubblicazione: (2026)
Monitoring Agentic Systems Before They're Reliable
di: Boston, Marisa Ferrara, et al.
Pubblicazione: (2026)
di: Boston, Marisa Ferrara, et al.
Pubblicazione: (2026)
AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code
di: Parris, William M.
Pubblicazione: (2026)
di: Parris, William M.
Pubblicazione: (2026)
Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
di: Baldonado, Juan Manuel, et al.
Pubblicazione: (2025)
di: Baldonado, Juan Manuel, et al.
Pubblicazione: (2025)
CodeTracer: Towards Traceable Agent States
di: Li, Han, et al.
Pubblicazione: (2026)
di: Li, Han, et al.
Pubblicazione: (2026)
LSPRAG: LSP-Guided RAG for Language-Agnostic Real-Time Unit Test Generation
di: Go, Gwihwan, et al.
Pubblicazione: (2025)
di: Go, Gwihwan, et al.
Pubblicazione: (2025)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
di: Chen, Jiachi, et al.
Pubblicazione: (2024)
di: Chen, Jiachi, et al.
Pubblicazione: (2024)
Kajal: Extracting Grammar of a Source Code Using Large Language Models
di: Torkamani, Mohammad Jalili
Pubblicazione: (2024)
di: Torkamani, Mohammad Jalili
Pubblicazione: (2024)
The Specification as Quality Gate: Three Hypotheses on AI-Assisted Code Review
di: Zietsman, Christo
Pubblicazione: (2026)
di: Zietsman, Christo
Pubblicazione: (2026)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
di: Ahmed, Sheikh Nazib, et al.
Pubblicazione: (2026)
di: Ahmed, Sheikh Nazib, et al.
Pubblicazione: (2026)
Automated structural testing of LLM-based agents: methods, framework, and case studies
di: Kohl, Jens, et al.
Pubblicazione: (2026)
di: Kohl, Jens, et al.
Pubblicazione: (2026)
Multi-Agent Code Verification via Information Theory
di: Rajan, Shreshth
Pubblicazione: (2025)
di: Rajan, Shreshth
Pubblicazione: (2025)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
di: Huang, Yuheng, et al.
Pubblicazione: (2024)
di: Huang, Yuheng, et al.
Pubblicazione: (2024)
Validating Solidity Code Defects using Symbolic and Concrete Execution powered by Large Language Models
di: Susan, Ştefan-Claudiu, et al.
Pubblicazione: (2025)
di: Susan, Ştefan-Claudiu, et al.
Pubblicazione: (2025)
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
di: Ravi, Ravin, et al.
Pubblicazione: (2026)
di: Ravi, Ravin, et al.
Pubblicazione: (2026)
Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution
di: Rao, Swanand
Pubblicazione: (2026)
di: Rao, Swanand
Pubblicazione: (2026)
CodeEvolve: LLM-Driven Evolutionary Optimization with Runtime-Enriched Target Selection for Multi-Language Code Enhancement
di: Borra, Ajay Krishna, et al.
Pubblicazione: (2026)
di: Borra, Ajay Krishna, et al.
Pubblicazione: (2026)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
di: Liu, Zhuoyao, et al.
Pubblicazione: (2026)
di: Liu, Zhuoyao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
di: Li, Meiziniu, et al.
Pubblicazione: (2022) -
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
di: Li, Meiziniu, et al.
Pubblicazione: (2024) -
Demystifying the Silence of Correctness Bugs in PyTorch Compiler
di: Li, Meiziniu, et al.
Pubblicazione: (2026) -
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
di: More, Riddhi, et al.
Pubblicazione: (2025) -
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
di: More, Riddhi, et al.
Pubblicazione: (2025)