LLMORPH: Automated Metamorphic Testing of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Steven, Ruberto, Stefano, Terragni, Valerio |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
by: Terragni, Valerio
Published: (2026)
by: Terragni, Valerio
Published: (2026)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
by: Huang, Yuheng, et al.
Published: (2024)
by: Huang, Yuheng, et al.
Published: (2024)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
by: Chang, Hung-Fu, et al.
Published: (2025)
by: Chang, Hung-Fu, et al.
Published: (2025)
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
by: Ravi, Ravin, et al.
Published: (2026)
by: Ravi, Ravin, et al.
Published: (2026)
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
by: More, Riddhi, et al.
Published: (2025)
by: More, Riddhi, et al.
Published: (2025)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
by: Weng, Haojun, et al.
Published: (2026)
by: Weng, Haojun, et al.
Published: (2026)
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
by: More, Riddhi, et al.
Published: (2025)
by: More, Riddhi, et al.
Published: (2025)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
by: Bradbury, Jeremy S., et al.
Published: (2024)
by: Bradbury, Jeremy S., et al.
Published: (2024)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
by: Chen, Jiachi, et al.
Published: (2024)
by: Chen, Jiachi, et al.
Published: (2024)
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
by: Jana, Prithwish, et al.
Published: (2023)
by: Jana, Prithwish, et al.
Published: (2023)
Automated structural testing of LLM-based agents: methods, framework, and case studies
by: Kohl, Jens, et al.
Published: (2026)
by: Kohl, Jens, et al.
Published: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
by: Baldonado, Juan Manuel, et al.
Published: (2025)
by: Baldonado, Juan Manuel, et al.
Published: (2025)
Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark
by: Mughal, Ali Hassaan, et al.
Published: (2026)
by: Mughal, Ali Hassaan, et al.
Published: (2026)
CIFE: Code Instruction-Following Evaluation
by: Gunnu, Sravani, et al.
Published: (2025)
by: Gunnu, Sravani, et al.
Published: (2025)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
by: Liu, Zhuoyao, et al.
Published: (2026)
by: Liu, Zhuoyao, et al.
Published: (2026)
Metamorphic Testing of Large Language Models for Natural Language Processing
by: Cho, Steven, et al.
Published: (2025)
by: Cho, Steven, et al.
Published: (2025)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG
by: Yu, Boxi, et al.
Published: (2026)
by: Yu, Boxi, et al.
Published: (2026)
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation
by: Holt, Samuel, et al.
Published: (2023)
by: Holt, Samuel, et al.
Published: (2023)
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
by: Mughal, Ali Hassaan, et al.
Published: (2026)
by: Mughal, Ali Hassaan, et al.
Published: (2026)
ASSERTIFY: Utilizing Large Language Models to Generate Assertions for Production Code
by: Torkamani, Mohammad Jalili, et al.
Published: (2024)
by: Torkamani, Mohammad Jalili, et al.
Published: (2024)
PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows
by: Kawamura, Kazuki, et al.
Published: (2026)
by: Kawamura, Kazuki, et al.
Published: (2026)
Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
by: Le, Nguyen-Khang, et al.
Published: (2025)
by: Le, Nguyen-Khang, et al.
Published: (2025)
Automated Bug Triaging using Instruction-Tuned Large Language Models
by: Kiashemshaki, Kiana, et al.
Published: (2025)
by: Kiashemshaki, Kiana, et al.
Published: (2025)
Leveraging Large Language Models for Use Case Model Generation from Software Requirements
by: Eisenreich, Tobias, et al.
Published: (2025)
by: Eisenreich, Tobias, et al.
Published: (2025)
RepoLaunch: Automating Build&Test Pipeline of Code Repositories on ANY Language and ANY Platform
by: Li, Kenan, et al.
Published: (2026)
by: Li, Kenan, et al.
Published: (2026)
Learning Software Bug Reports: A Systematic Literature Review
by: Long, Guoming, et al.
Published: (2025)
by: Long, Guoming, et al.
Published: (2025)
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
by: Li, Meiziniu, et al.
Published: (2022)
by: Li, Meiziniu, et al.
Published: (2022)
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems
by: Petrovic, Nenad, et al.
Published: (2024)
by: Petrovic, Nenad, et al.
Published: (2024)
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
by: Li, Meiziniu, et al.
Published: (2024)
by: Li, Meiziniu, et al.
Published: (2024)
Making a Pipeline Production-Ready: Challenges and Lessons Learned in the Healthcare Domain
by: Lawand, Daniel Angelo Esteves, et al.
Published: (2025)
by: Lawand, Daniel Angelo Esteves, et al.
Published: (2025)
SPIRA: Building an Intelligent System for Respiratory Insufficiency Detection
by: Ferreira, Renato Cordeiro, et al.
Published: (2025)
by: Ferreira, Renato Cordeiro, et al.
Published: (2025)
Distilling Desired Comments for Enhanced Code Review with Large Language Models
by: Yu, Yongda, et al.
Published: (2024)
by: Yu, Yongda, et al.
Published: (2024)
LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB
by: Bekmyradov, Vekil, et al.
Published: (2026)
by: Bekmyradov, Vekil, et al.
Published: (2026)
The Impact of Large Language Models on Open-source Innovation: Evidence from GitHub Copilot
by: Yeverechyahu, Doron, et al.
Published: (2024)
by: Yeverechyahu, Doron, et al.
Published: (2024)
Similar Items
-
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
by: Terragni, Valerio
Published: (2026) -
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
by: Huang, Yuheng, et al.
Published: (2024) -
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
by: Chang, Hung-Fu, et al.
Published: (2025) -
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
by: Ravi, Ravin, et al.
Published: (2026) -
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
by: More, Riddhi, et al.
Published: (2025)