MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Guoxiang, Aleti, Aldeida, Neelofar, Neelofar, Tantithamthavorn, Chakkrit, Qi, Yuanyuan, Chen, Tsong Yueh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PAFOT: A Position-Based Approach for Finding Optimal Tests of Autonomous Vehicles
by: Crespo-Rodriguez, Victor, et al.
Published: (2024)
by: Crespo-Rodriguez, Victor, et al.
Published: (2024)
UntrustVul: An Automated Approach for Identifying Untrustworthy Alerts in Vulnerability Detection Models
by: Tung, Lam Nguyen, et al.
Published: (2025)
by: Tung, Lam Nguyen, et al.
Published: (2025)
Automated Trustworthiness Oracle Generation for Machine Learning Text Classifiers
by: Tung, Lam Nguyen, et al.
Published: (2024)
by: Tung, Lam Nguyen, et al.
Published: (2024)
The Role of Road Features and Vehicle Dynamics in Cost-Effective Autonomous Vehicles Safety Testing: Insights from Instance Space Analysis
by: Crespo-Rodriguez, Victor, et al.
Published: (2026)
by: Crespo-Rodriguez, Victor, et al.
Published: (2026)
Requirements-Driven Automated Software Testing: A Systematic Review
by: Wang, Fanyu, et al.
Published: (2025)
by: Wang, Fanyu, et al.
Published: (2025)
From Domain Documents to Requirements: Retrieval-Augmented Generation in the Space Industry
by: Arora, Chetan, et al.
Published: (2025)
by: Arora, Chetan, et al.
Published: (2025)
Enhancing Large Language Models for Text-to-Testcase Generation
by: Alagarsamy, Saranya, et al.
Published: (2024)
by: Alagarsamy, Saranya, et al.
Published: (2024)
Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs
by: Wang, Fanyu, et al.
Published: (2025)
by: Wang, Fanyu, et al.
Published: (2025)
Neuron Patching: Semantic-based Neuron-level Language Model Repair for Code Generation
by: Gu, Jian, et al.
Published: (2023)
by: Gu, Jian, et al.
Published: (2023)
A Semantic-based Optimization Approach for Repairing LLMs: Case Study on Code Generation
by: Gu, Jian, et al.
Published: (2025)
by: Gu, Jian, et al.
Published: (2025)
Fine-Tuning and Prompt Engineering for Large Language Models-based Code Review Automation
by: Pornprasit, Chanathip, et al.
Published: (2024)
by: Pornprasit, Chanathip, et al.
Published: (2024)
Bidirectional Empowerment of Metamorphic Testing and Large Language Models: A Systematic Survey
by: Zheng, Zheng, et al.
Published: (2026)
by: Zheng, Zheng, et al.
Published: (2026)
Code Ownership: The Principles, Differences, and Their Associations with Software Quality
by: Thongtanunam, Patanamon, et al.
Published: (2024)
by: Thongtanunam, Patanamon, et al.
Published: (2024)
Code Readability in the Age of Large Language Models: An Industrial Case Study from Atlassian
by: Takerngsaksiri, Wannita, et al.
Published: (2025)
by: Takerngsaksiri, Wannita, et al.
Published: (2025)
Experimental evaluation of architectural software performance design patterns in microservices
by: Meijer, Willem, et al.
Published: (2024)
by: Meijer, Willem, et al.
Published: (2024)
Test-based Patch Clustering for Automatically-Generated Patches Assessment
by: Martinez, Matias, et al.
Published: (2022)
by: Martinez, Matias, et al.
Published: (2022)
Trustworthy AI Software Engineers
by: Aleti, Aldeida, et al.
Published: (2026)
by: Aleti, Aldeida, et al.
Published: (2026)
Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
by: Salimian, Sina, et al.
Published: (2025)
by: Salimian, Sina, et al.
Published: (2025)
Metamorphic Testing for Audio Content Moderation Software
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
AI for DevSecOps: A Landscape and Future Opportunities
by: Fu, Michael, et al.
Published: (2024)
by: Fu, Michael, et al.
Published: (2024)
Navigating Fairness: Practitioners' Understanding, Challenges, and Strategies in AI/ML Development
by: Pant, Aastha, et al.
Published: (2024)
by: Pant, Aastha, et al.
Published: (2024)
LLM-Based Static Verification of Code Against Natural-Language Requirements: An Industrial Experience Report
by: Zhou, Zhi Quan, et al.
Published: (2026)
by: Zhou, Zhi Quan, et al.
Published: (2026)
On the Reliability and Explainability of Language Models for Program Generation
by: Liu, Yue, et al.
Published: (2023)
by: Liu, Yue, et al.
Published: (2023)
Ethics in AI through the Practitioner's View: A Grounded Theory Literature Review
by: Pant, Aastha, et al.
Published: (2022)
by: Pant, Aastha, et al.
Published: (2022)
What do AI/ML practitioners think about AI/ML bias?
by: Pant, Aastha, et al.
Published: (2024)
by: Pant, Aastha, et al.
Published: (2024)
ViBR: Automated Bug Replay from Video-based Reports using Vision-Language Models
by: Feng, Sidong, et al.
Published: (2026)
by: Feng, Sidong, et al.
Published: (2026)
Metamorphic Relation Generation: State of the Art and Visions for Future Research
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
From Particles to Perils: SVGD-Based Hazardous Scenario Generation for Autonomous Driving Systems Testing
by: Liang, Linfeng, et al.
Published: (2026)
by: Liang, Linfeng, et al.
Published: (2026)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
PyTester: Deep Reinforcement Learning for Text-to-Testcase Generation
by: Takerngsaksiri, Wannita, et al.
Published: (2024)
by: Takerngsaksiri, Wannita, et al.
Published: (2024)
Integrating Artificial Intelligence with Human Expertise: An In-depth Analysis of ChatGPT's Capabilities in Generating Metamorphic Relations
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Protect Your Secrets: Understanding and Measuring Data Exposure in VSCode Extensions
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Efficient Fairness Testing in Large Language Models: Prioritizing Metamorphic Relations for Bias Detection
by: Giramata, Suavis, et al.
Published: (2025)
by: Giramata, Suavis, et al.
Published: (2025)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
by: Wang, Sizhe, et al.
Published: (2025)
by: Wang, Sizhe, et al.
Published: (2025)
Metamorphic Debugging for Accountable Software
by: Tizpaz-Niari, Saeid, et al.
Published: (2024)
by: Tizpaz-Niari, Saeid, et al.
Published: (2024)
Software-Based Dialogue Systems: Survey, Taxonomy and Challenges
by: Motger, Quim, et al.
Published: (2021)
by: Motger, Quim, et al.
Published: (2021)
RAGVA: Engineering Retrieval Augmented Generation-based Virtual Assistants in Practice
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Blended PC Peer Review Model: Process and Reflection
by: Tantithamthavorn, Chakkrit, et al.
Published: (2025)
by: Tantithamthavorn, Chakkrit, et al.
Published: (2025)
Automatically Recommend Code Updates: Are We There Yet?
by: Liu, Yue, et al.
Published: (2022)
by: Liu, Yue, et al.
Published: (2022)
On the Potential and Limitations of Few-Shot In-Context Learning to Generate Metamorphic Specifications for Tax Preparation Software
by: Srinivas, Dananjay, et al.
Published: (2023)
by: Srinivas, Dananjay, et al.
Published: (2023)
Similar Items
-
PAFOT: A Position-Based Approach for Finding Optimal Tests of Autonomous Vehicles
by: Crespo-Rodriguez, Victor, et al.
Published: (2024) -
UntrustVul: An Automated Approach for Identifying Untrustworthy Alerts in Vulnerability Detection Models
by: Tung, Lam Nguyen, et al.
Published: (2025) -
Automated Trustworthiness Oracle Generation for Machine Learning Text Classifiers
by: Tung, Lam Nguyen, et al.
Published: (2024) -
The Role of Road Features and Vehicle Dynamics in Cost-Effective Autonomous Vehicles Safety Testing: Insights from Instance Space Analysis
by: Crespo-Rodriguez, Victor, et al.
Published: (2026) -
Requirements-Driven Automated Software Testing: A Systematic Review
by: Wang, Fanyu, et al.
Published: (2025)