Private GPTs for LLM-driven testing in software development and machine learning
Fuente:
arXiv
Saved in:
| Main Authors: | Jagielski, Jakub, Rojas, Consuelo, Abel, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fuzzing the brain: Automated stress testing for the safety of ML-driven neurostimulation
by: Downing, Mara, et al.
Published: (2025)
by: Downing, Mara, et al.
Published: (2025)
FREYR: A Framework for Recognizing and Executing Your Requests
by: Gallotta, Roberto, et al.
Published: (2025)
by: Gallotta, Roberto, et al.
Published: (2025)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
by: Karpurapu, Shanthi, et al.
Published: (2024)
by: Karpurapu, Shanthi, et al.
Published: (2024)
SLEGO: A Collaborative Data Analytics System with LLM Recommender for Diverse Users
by: Ng, Siu Lung, et al.
Published: (2024)
by: Ng, Siu Lung, et al.
Published: (2024)
NormCode Canvas: Making LLM Agentic Workflows Development Sustainable via Case-Based Reasoning
by: Guan, Xin, et al.
Published: (2026)
by: Guan, Xin, et al.
Published: (2026)
Test-driven Software Experimentation with LASSO: an LLM Prompt Benchmarking Example
by: Kessel, Marcus
Published: (2024)
by: Kessel, Marcus
Published: (2024)
Social, Legal, Ethical, Empathetic and Cultural Norm Operationalisation for AI Agents
by: Calinescu, Radu, et al.
Published: (2026)
by: Calinescu, Radu, et al.
Published: (2026)
LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB
by: Bekmyradov, Vekil, et al.
Published: (2026)
by: Bekmyradov, Vekil, et al.
Published: (2026)
Smart Expansion Techniques for ASP-based Interactive Configuration
by: Balážová, Lucia, et al.
Published: (2025)
by: Balážová, Lucia, et al.
Published: (2025)
AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
by: Bhardwaj, Varun Pratap
Published: (2026)
by: Bhardwaj, Varun Pratap
Published: (2026)
Inference-Time Intervention in Large Language Models for Reliable Requirement Verification
by: Darm, Paul, et al.
Published: (2025)
by: Darm, Paul, et al.
Published: (2025)
From Internet of Things Data to Business Processes: Challenges and a Framework
by: Mangler, Juergen, et al.
Published: (2024)
by: Mangler, Juergen, et al.
Published: (2024)
Who is Introducing the Failure? Automatically Attributing Failures of Multi-Agent Systems via Spectrum Analysis
by: Ge, Yu, et al.
Published: (2025)
by: Ge, Yu, et al.
Published: (2025)
Memory Management and Contextual Consistency for Long-Running Low-Code Agents
by: Xu, Jiexi
Published: (2025)
by: Xu, Jiexi
Published: (2025)
Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework
by: Zietsman, Christo
Published: (2026)
by: Zietsman, Christo
Published: (2026)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
by: Rehan, Tzafrir
Published: (2026)
by: Rehan, Tzafrir
Published: (2026)
N-Version Assessment and Enhancement of Generative AI
by: Kessel, Marcus, et al.
Published: (2024)
by: Kessel, Marcus, et al.
Published: (2024)
Morescient GAI for Software Engineering (Extended Version)
by: Kessel, Marcus, et al.
Published: (2024)
by: Kessel, Marcus, et al.
Published: (2024)
AI-Assisted Engineering Should Track the Epistemic Status and Temporal Validity of Architectural Decisions
by: Gilda, Sankalp, et al.
Published: (2026)
by: Gilda, Sankalp, et al.
Published: (2026)
An Industrial-Scale Retrieval-Augmented Generation Framework for Requirements Engineering: Empirical Evaluation with Automotive Manufacturing Data
by: Khalid, Muhammad, et al.
Published: (2026)
by: Khalid, Muhammad, et al.
Published: (2026)
Bounded Autonomy for Enterprise AI: Typed Action Contracts and Consumer-Side Execution
by: Sohail, Sarmad, et al.
Published: (2026)
by: Sohail, Sarmad, et al.
Published: (2026)
From Machine Learning Documentation to Requirements: Bridging Processes with Requirements Languages
by: Peng, Yi, et al.
Published: (2025)
by: Peng, Yi, et al.
Published: (2025)
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
by: Rodriguez, David, et al.
Published: (2025)
by: Rodriguez, David, et al.
Published: (2025)
Reasoning Provenance for Autonomous AI Agents: Structured Behavioral Analytics Beyond State Checkpoints and Execution Traces
by: Vispute, Neelmani, et al.
Published: (2026)
by: Vispute, Neelmani, et al.
Published: (2026)
Exploring LLMs for User Story Extraction from Mockups
by: Firmenich, Diego, et al.
Published: (2026)
by: Firmenich, Diego, et al.
Published: (2026)
AutoReSpec: A Framework for Generating Specification using Large Language Models
by: Ayon, Ragib Shahariar, et al.
Published: (2026)
by: Ayon, Ragib Shahariar, et al.
Published: (2026)
Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
Evaluating LLM-driven User-Intent Formalization for Verification-Aware Languages
by: Lahiri, Shuvendu K.
Published: (2024)
by: Lahiri, Shuvendu K.
Published: (2024)
A domain-specific language for describing machine learning datasets
by: Giner-Miguelez, Joan, et al.
Published: (2022)
by: Giner-Miguelez, Joan, et al.
Published: (2022)
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
by: Kohl, Jens, et al.
Published: (2024)
by: Kohl, Jens, et al.
Published: (2024)
IFRA: a machine learning-based Instrumented Fall Risk Assessment Scale derived from Instrumented Timed Up and Go test in stroke patients
by: Macciò, Simone, et al.
Published: (2025)
by: Macciò, Simone, et al.
Published: (2025)
CodeEvolve: LLM-Driven Evolutionary Optimization with Runtime-Enriched Target Selection for Multi-Language Code Enhancement
by: Borra, Ajay Krishna, et al.
Published: (2026)
by: Borra, Ajay Krishna, et al.
Published: (2026)
ContractBench: Can LLM Agents Preserve Observation Contracts?
by: Wang, Jicheng, et al.
Published: (2026)
by: Wang, Jicheng, et al.
Published: (2026)
A Framework for Testing and Adapting REST APIs as LLM Tools
by: Bandlamudi, Jayachandu, et al.
Published: (2025)
by: Bandlamudi, Jayachandu, et al.
Published: (2025)
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems
by: Petrovic, Nenad, et al.
Published: (2024)
by: Petrovic, Nenad, et al.
Published: (2024)
Towards Single-System Illusion in Software-Defined Vehicles -- Automated, AI-Powered Workflow
by: Lebioda, Krzysztof, et al.
Published: (2024)
by: Lebioda, Krzysztof, et al.
Published: (2024)
Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
by: Cartagena, Arnold, et al.
Published: (2026)
by: Cartagena, Arnold, et al.
Published: (2026)
Large Language Models for Combinatorial Optimization of Design Structure Matrix
by: Jiang, Shuo, et al.
Published: (2025)
by: Jiang, Shuo, et al.
Published: (2025)
Intelligent Design 4.0: Paradigm Evolution Toward the Agentic AI Era
by: Jiang, Shuo, et al.
Published: (2025)
by: Jiang, Shuo, et al.
Published: (2025)
Design Structure Matrix Modularization with Large Language Models
by: Jiang, Shuo, et al.
Published: (2026)
by: Jiang, Shuo, et al.
Published: (2026)
Similar Items
-
Fuzzing the brain: Automated stress testing for the safety of ML-driven neurostimulation
by: Downing, Mara, et al.
Published: (2025) -
FREYR: A Framework for Recognizing and Executing Your Requests
by: Gallotta, Roberto, et al.
Published: (2025) -
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
by: Karpurapu, Shanthi, et al.
Published: (2024) -
SLEGO: A Collaborative Data Analytics System with LLM Recommender for Diverse Users
by: Ng, Siu Lung, et al.
Published: (2024) -
NormCode Canvas: Making LLM Agentic Workflows Development Sustainable via Case-Based Reasoning
by: Guan, Xin, et al.
Published: (2026)