Towards Test Generation from Task Description for Mobile Testing with Multi-modal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Huynh, Hieu, Phung, Hai, Pham, Hao, Nguyen, Tien N., Nguyen, Vu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Segment-Based Test Case Prioritization: A Multi-objective Approach
by: Huynh, Hieu, et al.
Published: (2024)
by: Huynh, Hieu, et al.
Published: (2024)
Toward Generation of Test Cases from Task Descriptions via History-aware Planning
by: Cao, Duy, et al.
Published: (2025)
by: Cao, Duy, et al.
Published: (2025)
RBCTest: Leveraging LLMs to Mine and Verify Oracles of API Response Bodies for RESTful API Testing
by: Huynh, Hieu, et al.
Published: (2025)
by: Huynh, Hieu, et al.
Published: (2025)
Reinforcement Learning-Based REST API Testing with Multi-Coverage
by: Nguyen, Tien-Quang, et al.
Published: (2024)
by: Nguyen, Tien-Quang, et al.
Published: (2024)
TestWeaver: Execution-aware, Feedback-driven Regression Testing Generation with Large Language Models
by: Le, Cuong Chi, et al.
Published: (2025)
by: Le, Cuong Chi, et al.
Published: (2025)
KAT: Dependency-aware Automated API Testing with Large Language Models
by: Le, Tri, et al.
Published: (2024)
by: Le, Tri, et al.
Published: (2024)
Automated Description Generation for Software Patches
by: Vu, Thanh Trong, et al.
Published: (2024)
by: Vu, Thanh Trong, et al.
Published: (2024)
Generating Critical Scenarios for Testing Automated Driving Systems
by: Nguyen, Trung-Hieu, et al.
Published: (2024)
by: Nguyen, Trung-Hieu, et al.
Published: (2024)
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
by: Manh, Dung Nguyen, et al.
Published: (2024)
by: Manh, Dung Nguyen, et al.
Published: (2024)
Model-Agnostic Correctness Assessment for LLM-Generated Code via Dynamic Internal Representation Selection
by: Vu, Thanh Trong, et al.
Published: (2025)
by: Vu, Thanh Trong, et al.
Published: (2025)
CABENCH: Benchmarking Composable AI for Solving Complex Tasks through Composing Ready-to-Use Models
by: Pham, Tung-Thuy, et al.
Published: (2025)
by: Pham, Tung-Thuy, et al.
Published: (2025)
Cerberus: Multi-Agent Reasoning and Coverage-Guided Exploration for Static Detection of Runtime Errors
by: Dhulipala, Hridya, et al.
Published: (2025)
by: Dhulipala, Hridya, et al.
Published: (2025)
Towards Multi-Platform Mutation Testing of Task-based Chatbots
by: Clerissi, Diego, et al.
Published: (2025)
by: Clerissi, Diego, et al.
Published: (2025)
Sakura: An Approach for Generating Complex Tests from Natural Language Test Descriptions
by: Stennett, Tyler, et al.
Published: (2026)
by: Stennett, Tyler, et al.
Published: (2026)
When Retriever Meets Generator: A Joint Model for Code Comment Generation
by: Le, Tien P. T., et al.
Published: (2025)
by: Le, Tien P. T., et al.
Published: (2025)
Generating REST API Tests With Descriptive Names
by: Garrett, Philip, et al.
Published: (2025)
by: Garrett, Philip, et al.
Published: (2025)
Correctness Assessment of Code Generated by Large Language Models Using Internal Representations
by: Bui, Tuan-Dung, et al.
Published: (2025)
by: Bui, Tuan-Dung, et al.
Published: (2025)
An Empirical Study on Capability of Large Language Models in Understanding Code Semantics
by: Nguyen, Thu-Trang, et al.
Published: (2024)
by: Nguyen, Thu-Trang, et al.
Published: (2024)
MEMRES: A Memory-Augmented Resolver with Confidence Cascade for Agentic Python Dependency Resolution
by: Minh, Dao Sy Duy, et al.
Published: (2026)
by: Minh, Dao Sy Duy, et al.
Published: (2026)
CodeLSI: Leveraging Foundation Models for Automated Code Generation with Low-Rank Optimization and Domain-Specific Instruction Tuning
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
SpecMind: Cognitively Inspired, Interactive Multi-Turn Framework for Postcondition Inference
by: Le, Cuong Chi, et al.
Published: (2026)
by: Le, Cuong Chi, et al.
Published: (2026)
HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale
by: Phan, Huy Nhat, et al.
Published: (2024)
by: Phan, Huy Nhat, et al.
Published: (2024)
The Effect of Code Obfuscation on Human Program Comprehension
by: Nguyen, Anh H. N., et al.
Published: (2026)
by: Nguyen, Anh H. N., et al.
Published: (2026)
Toward Realistic Evaluations of Just-In-Time Vulnerability Prediction
by: Nguyen, Duong, et al.
Published: (2025)
by: Nguyen, Duong, et al.
Published: (2025)
Scaling Mobile Chaos Testing with AI-Driven Test Execution
by: Marcano, Juan, et al.
Published: (2026)
by: Marcano, Juan, et al.
Published: (2026)
Testing Is Not Boring: Characterizing Challenge in Software Testing Tasks
by: Hardman, Davi Gama, et al.
Published: (2025)
by: Hardman, Davi Gama, et al.
Published: (2025)
Test Case Generation for Dialogflow Task-Based Chatbots
by: Rapisarda, Rocco Gianni, et al.
Published: (2025)
by: Rapisarda, Rocco Gianni, et al.
Published: (2025)
Program Structure-aware Language Models: Targeted Software Testing beyond Textual Semantics
by: Tran, Khang, et al.
Published: (2026)
by: Tran, Khang, et al.
Published: (2026)
Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing
by: Tu, Zhi, et al.
Published: (2025)
by: Tu, Zhi, et al.
Published: (2025)
Automated Functional Testing for Malleable Mobile Application Driven from User Intent
by: Wang, Yuying, et al.
Published: (2026)
by: Wang, Yuying, et al.
Published: (2026)
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
by: Le-Cong, Thanh, et al.
Published: (2024)
by: Le-Cong, Thanh, et al.
Published: (2024)
FuncDroid: Towards Inter-Functional Flows for Comprehensive Mobile App GUI Testing
by: He, Jinlong, et al.
Published: (2026)
by: He, Jinlong, et al.
Published: (2026)
Verifying DNN-based Semantic Communication Against Generative Adversarial Noise
by: Le, Thanh, et al.
Published: (2026)
by: Le, Thanh, et al.
Published: (2026)
PrismaDV: Automated Task-Aware Data Unit Test Generation
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
Fuzzwise: Intelligent Initial Corpus Generation for Fuzzing
by: Dhulipala, Hridya, et al.
Published: (2025)
by: Dhulipala, Hridya, et al.
Published: (2025)
Data-Driven Evidence-Based Syntactic Sugar Design
by: OBrien, David, et al.
Published: (2024)
by: OBrien, David, et al.
Published: (2024)
Enhancing Program Repair with Specification Guidance and Intermediate Behavioral Signals
by: Le-Anh, Minh, et al.
Published: (2026)
by: Le-Anh, Minh, et al.
Published: (2026)
From Exploration to Specification: LLM-Based Property Generation for Mobile App Testing
by: Xiong, Yiheng, et al.
Published: (2026)
by: Xiong, Yiheng, et al.
Published: (2026)
Polynomiogram: An Integrated Framework for Root Visualization and Generative Art
by: Nguyen, Hoang Duc, et al.
Published: (2025)
by: Nguyen, Hoang Duc, et al.
Published: (2025)
Are the Majority of Public Computational Notebooks Pathologically Non-Executable?
by: Nguyen, Tien, et al.
Published: (2025)
by: Nguyen, Tien, et al.
Published: (2025)
Similar Items
-
Segment-Based Test Case Prioritization: A Multi-objective Approach
by: Huynh, Hieu, et al.
Published: (2024) -
Toward Generation of Test Cases from Task Descriptions via History-aware Planning
by: Cao, Duy, et al.
Published: (2025) -
RBCTest: Leveraging LLMs to Mine and Verify Oracles of API Response Bodies for RESTful API Testing
by: Huynh, Hieu, et al.
Published: (2025) -
Reinforcement Learning-Based REST API Testing with Multi-Coverage
by: Nguyen, Tien-Quang, et al.
Published: (2024) -
TestWeaver: Execution-aware, Feedback-driven Regression Testing Generation with Large Language Models
by: Le, Cuong Chi, et al.
Published: (2025)