Automated Testing of Task-based Chatbots: How Far Are We?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Clerissi, Diego, Masserini, Elena, Micucci, Daniela, Mariani, Leonardo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Multi-Platform Mutation Testing of Task-based Chatbots
von: Clerissi, Diego, et al.
Veröffentlicht: (2025)
von: Clerissi, Diego, et al.
Veröffentlicht: (2025)
Assessing Task-based Chatbots: Snapshot and Curated Datasets for Dialogflow
von: Masserini, Elena, et al.
Veröffentlicht: (2026)
von: Masserini, Elena, et al.
Veröffentlicht: (2026)
Towards the Assessment of Task-based Chatbots: From the TOFU-R Snapshot to the BRASATO Curated Dataset
von: Masserini, Elena, et al.
Veröffentlicht: (2025)
von: Masserini, Elena, et al.
Veröffentlicht: (2025)
Bug Whispering: Towards Audio Bug Reporting
von: Masserini, Elena, et al.
Veröffentlicht: (2025)
von: Masserini, Elena, et al.
Veröffentlicht: (2025)
Anonymizing Test Data in Android: Does It Hurt?
von: Masserini, Elena, et al.
Veröffentlicht: (2024)
von: Masserini, Elena, et al.
Veröffentlicht: (2024)
Test Case Generation for Dialogflow Task-Based Chatbots
von: Rapisarda, Rocco Gianni, et al.
Veröffentlicht: (2025)
von: Rapisarda, Rocco Gianni, et al.
Veröffentlicht: (2025)
MutaBot: A Mutation Testing Approach for Chatbots
von: Urrico, Michael Ferdinando, et al.
Veröffentlicht: (2024)
von: Urrico, Michael Ferdinando, et al.
Veröffentlicht: (2024)
Assessing AI-Based Code Assistants in Method Generation Tasks
von: Corso, Vincenzo, et al.
Veröffentlicht: (2024)
von: Corso, Vincenzo, et al.
Veröffentlicht: (2024)
Studying How Configurations Impact Code Generation in LLMs: the Case of ChatGPT
von: Donato, Benedetta, et al.
Veröffentlicht: (2025)
von: Donato, Benedetta, et al.
Veröffentlicht: (2025)
Multi-Level Testing of Conversational AI Systems
von: Masserini, Elena
Veröffentlicht: (2026)
von: Masserini, Elena
Veröffentlicht: (2026)
MultiMind: A Plug-in for the Implementation of Development Tasks Aided by AI Assistants
von: Donato, Benedetta, et al.
Veröffentlicht: (2025)
von: Donato, Benedetta, et al.
Veröffentlicht: (2025)
Analyzing Prompt Influence on Automated Method Generation: An Empirical Study with Copilot
von: Fagadau, Ionut Daniel, et al.
Veröffentlicht: (2024)
von: Fagadau, Ionut Daniel, et al.
Veröffentlicht: (2024)
Generating Java Methods: An Empirical Assessment of Four AI-Based Code Assistants
von: Corso, Vincenzo, et al.
Veröffentlicht: (2024)
von: Corso, Vincenzo, et al.
Veröffentlicht: (2024)
On the Possibility of Breaking Copyleft Licenses When Reusing Code Generated by ChatGPT
von: Colombo, Gaia, et al.
Veröffentlicht: (2025)
von: Colombo, Gaia, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Test Generation: How Far Are We?
von: Shin, Jiho, et al.
Veröffentlicht: (2024)
von: Shin, Jiho, et al.
Veröffentlicht: (2024)
Deep Learning Framework Testing via Model Mutation: How Far Are We?
von: Mu, Yanzhou, et al.
Veröffentlicht: (2025)
von: Mu, Yanzhou, et al.
Veröffentlicht: (2025)
Static Application Security Testing (SAST) Tools for Smart Contracts: How Far Are We?
von: Li, Kaixuan, et al.
Veröffentlicht: (2024)
von: Li, Kaixuan, et al.
Veröffentlicht: (2024)
Duplicate Bug Report Detection: How Far Are We?
von: Zhang, Ting, et al.
Veröffentlicht: (2022)
von: Zhang, Ting, et al.
Veröffentlicht: (2022)
Vulnerability-Affected Versions Identification: How Far Are We?
von: Chen, Xingchu, et al.
Veröffentlicht: (2025)
von: Chen, Xingchu, et al.
Veröffentlicht: (2025)
TESTQUEST: A Web Gamification Tool to Improve Locators and Page Objects Quality
von: Olianas, Dario, et al.
Veröffentlicht: (2025)
von: Olianas, Dario, et al.
Veröffentlicht: (2025)
Root Cause Analysis for Microservice System based on Causal Inference: How Far Are We?
von: Pham, Luan, et al.
Veröffentlicht: (2024)
von: Pham, Luan, et al.
Veröffentlicht: (2024)
Can AI Agents Generate Microservices? How Far are We?
von: Adnan, Bassam, et al.
Veröffentlicht: (2026)
von: Adnan, Bassam, et al.
Veröffentlicht: (2026)
Representation Learning for Stack Overflow Posts: How Far are We?
von: He, Junda, et al.
Veröffentlicht: (2023)
von: He, Junda, et al.
Veröffentlicht: (2023)
Large Language Models for Equivalent Mutant Detection: How Far Are We?
von: Tian, Zhao, et al.
Veröffentlicht: (2024)
von: Tian, Zhao, et al.
Veröffentlicht: (2024)
Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We?
von: O'Brien, Conor, et al.
Veröffentlicht: (2024)
von: O'Brien, Conor, et al.
Veröffentlicht: (2024)
Unraveling the Potential of Large Language Models in Code Translation: How Far Are We?
von: Tao, Qingxiao, et al.
Veröffentlicht: (2024)
von: Tao, Qingxiao, et al.
Veröffentlicht: (2024)
A Large-Scale Evaluation for Log Parsing Techniques: How Far Are We?
von: Jiang, Zhihan, et al.
Veröffentlicht: (2023)
von: Jiang, Zhihan, et al.
Veröffentlicht: (2023)
How Far Can We Go with Practical Function-Level Program Repair?
von: Xiang, Jiahong, et al.
Veröffentlicht: (2024)
von: Xiang, Jiahong, et al.
Veröffentlicht: (2024)
Leveraging Language Models for Log Statement Generation in Multilingual Scenarios: How Far Are We?
von: Kusama, Kazuki, et al.
Veröffentlicht: (2026)
von: Kusama, Kazuki, et al.
Veröffentlicht: (2026)
When ChatGPT Meets Smart Contract Vulnerability Detection: How Far Are We?
von: Chen, Chong, et al.
Veröffentlicht: (2023)
von: Chen, Chong, et al.
Veröffentlicht: (2023)
Specification-Driven Code Translation Powered by Large Language Models: How Far Are We?
von: Saha, Soumit Kanti, et al.
Veröffentlicht: (2024)
von: Saha, Soumit Kanti, et al.
Veröffentlicht: (2024)
An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far Are We?
von: Suh, Hyunjae, et al.
Veröffentlicht: (2024)
von: Suh, Hyunjae, et al.
Veröffentlicht: (2024)
Vulnerability Detection with Code Language Models: How Far Are We?
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
Model Editing for LLMs4Code: How Far are We?
von: Li, Xiaopeng, et al.
Veröffentlicht: (2024)
von: Li, Xiaopeng, et al.
Veröffentlicht: (2024)
Using LLMs for Security Advisory Investigations: How Far Are We?
von: Abdullah, Bayu Fedra, et al.
Veröffentlicht: (2025)
von: Abdullah, Bayu Fedra, et al.
Veröffentlicht: (2025)
LLM For Loop Invariant Generation and Fixing: How Far Are We?
von: Akhond, Mostafijur Rahman, et al.
Veröffentlicht: (2025)
von: Akhond, Mostafijur Rahman, et al.
Veröffentlicht: (2025)
Testing in the Evolving World of DL Systems:Insights from Python GitHub Projects
von: Ali, Qurban, et al.
Veröffentlicht: (2024)
von: Ali, Qurban, et al.
Veröffentlicht: (2024)
How Far Have LLMs Come Toward Automated SATD Taxonomy Construction?
von: Nakashima, Sota, et al.
Veröffentlicht: (2025)
von: Nakashima, Sota, et al.
Veröffentlicht: (2025)
OpenCat: Improving Interoperability of ADS Testing
von: Ali, Qurban, et al.
Veröffentlicht: (2025)
von: Ali, Qurban, et al.
Veröffentlicht: (2025)
Reasoning Runtime Behavior of a Program with LLM: How Far Are We?
von: Chen, Junkai, et al.
Veröffentlicht: (2024)
von: Chen, Junkai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Multi-Platform Mutation Testing of Task-based Chatbots
von: Clerissi, Diego, et al.
Veröffentlicht: (2025) -
Assessing Task-based Chatbots: Snapshot and Curated Datasets for Dialogflow
von: Masserini, Elena, et al.
Veröffentlicht: (2026) -
Towards the Assessment of Task-based Chatbots: From the TOFU-R Snapshot to the BRASATO Curated Dataset
von: Masserini, Elena, et al.
Veröffentlicht: (2025) -
Bug Whispering: Towards Audio Bug Reporting
von: Masserini, Elena, et al.
Veröffentlicht: (2025) -
Anonymizing Test Data in Android: Does It Hurt?
von: Masserini, Elena, et al.
Veröffentlicht: (2024)