Saved in:
| Main Authors: | Perera, Wanni Vidulige Ishan, Liu, Xing, liang, Fan, Zhang, Junyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.07933 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoDroid: LLM-powered Task Automation in Android
by: Wen, Hao, et al.
Published: (2023)
by: Wen, Hao, et al.
Published: (2023)
Breaking Barriers in Software Testing: The Power of AI-Driven Automation
by: Naqvi, Saba, et al.
Published: (2025)
by: Naqvi, Saba, et al.
Published: (2025)
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
by: Son, Ha Min, et al.
Published: (2025)
by: Son, Ha Min, et al.
Published: (2025)
On the Adoption of AI Coding Agents in Open-source Android and iOS Development
by: Khan, Muhammad Ahmad, et al.
Published: (2026)
by: Khan, Muhammad Ahmad, et al.
Published: (2026)
Understanding the Weakness of Large Language Model Agents within a Complex Android Environment
by: Xing, Mingzhe, et al.
Published: (2024)
by: Xing, Mingzhe, et al.
Published: (2024)
Breaking the Illusion of Identity in LLM Tooling
by: Miller, Marek
Published: (2026)
by: Miller, Marek
Published: (2026)
DroidBot-GPT: GPT-powered UI Automation for Android
by: Wen, Hao, et al.
Published: (2023)
by: Wen, Hao, et al.
Published: (2023)
A Deep Dive Into Large Language Model Code Generation Mistakes: What and Why?
by: Chen, QiHong, et al.
Published: (2024)
by: Chen, QiHong, et al.
Published: (2024)
Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench
by: Zou, Qingyun, et al.
Published: (2026)
by: Zou, Qingyun, et al.
Published: (2026)
DynamicsLLM: a Dynamic Analysis-based Tool for Generating Intelligent Execution Traces Using LLMs to Detect Android Behavioural Code Smells
by: Cherief, Houcine Abdelkader, et al.
Published: (2026)
by: Cherief, Houcine Abdelkader, et al.
Published: (2026)
MASKDROID: Robust Android Malware Detection with Masked Graph Representations
by: Zheng, Jingnan, et al.
Published: (2024)
by: Zheng, Jingnan, et al.
Published: (2024)
Small Models, Big Tasks: An Exploratory Empirical Study on Small Language Models for Function Calling
by: Kavathekar, Ishan, et al.
Published: (2025)
by: Kavathekar, Ishan, et al.
Published: (2025)
AndroidControl-Curated: Revealing the True Potential of GUI Agents through Benchmark Purification
by: Leung, Ho Fai, et al.
Published: (2025)
by: Leung, Ho Fai, et al.
Published: (2025)
CodeFuse-CommitEval: Towards Benchmarking LLM's Power on Commit Message and Code Change Inconsistency Detection
by: Zhang, Qingyu, et al.
Published: (2025)
by: Zhang, Qingyu, et al.
Published: (2025)
AI-Powered Commit Explorer (APCE)
by: Grees, Yousab, et al.
Published: (2025)
by: Grees, Yousab, et al.
Published: (2025)
exLong: Generating Exceptional Behavior Tests with Large Language Models
by: Zhang, Jiyang, et al.
Published: (2024)
by: Zhang, Jiyang, et al.
Published: (2024)
SELF-REDRAFT: Eliciting Intrinsic Exploration-Exploitation Balance in Test-Time Scaling for Code Generation
by: Chen, Yixiang, et al.
Published: (2025)
by: Chen, Yixiang, et al.
Published: (2025)
Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
by: Rombaut, Benjamin, et al.
Published: (2024)
by: Rombaut, Benjamin, et al.
Published: (2024)
Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
by: Xia, Boming, et al.
Published: (2024)
by: Xia, Boming, et al.
Published: (2024)
A Tool for Generating Exceptional Behavior Tests With Large Language Models
by: Zhong, Linghan, et al.
Published: (2025)
by: Zhong, Linghan, et al.
Published: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
Breaking the Myth: Can Small Models Infer Postconditions Too?
by: Zhang, Gehao, et al.
Published: (2025)
by: Zhang, Gehao, et al.
Published: (2025)
AI2Apps: A Visual IDE for Building LLM-based AI Agent Applications
by: Pang, Xin, et al.
Published: (2024)
by: Pang, Xin, et al.
Published: (2024)
LLM-Powered Code Vulnerability Repair with Reinforcement Learning and Semantic Reward
by: Islam, Nafis Tanveer, et al.
Published: (2024)
by: Islam, Nafis Tanveer, et al.
Published: (2024)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
by: Pan, Zhiyuan, et al.
Published: (2025)
by: Pan, Zhiyuan, et al.
Published: (2025)
KGCompiler: Deep Learning Compilation Optimization for Knowledge Graph Complex Logical Query Answering
by: Lin, Hongyu, et al.
Published: (2025)
by: Lin, Hongyu, et al.
Published: (2025)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)
by: Peng, Jiaren, et al.
Published: (2026)
Still Manual? Automated Linter Configuration via DSL-Based LLM Compilation of Coding Standards
by: Zhang, Zejun, et al.
Published: (2026)
by: Zhang, Zejun, et al.
Published: (2026)
iPanda: An LLM-based Agent for Automated Conformance Testing of Communication Protocols
by: Sun, Xikai, et al.
Published: (2025)
by: Sun, Xikai, et al.
Published: (2025)
Self-Healing Software Systems: Lessons from Nature, Powered by AI
by: Baqar, Mohammad, et al.
Published: (2025)
by: Baqar, Mohammad, et al.
Published: (2025)
The Future of Software Testing: AI-Powered Test Case Generation and Validation
by: Baqar, Mohammad, et al.
Published: (2024)
by: Baqar, Mohammad, et al.
Published: (2024)
When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs'' for Human-AI Interaction
by: Xing, Zhenchang, et al.
Published: (2025)
by: Xing, Zhenchang, et al.
Published: (2025)
LLM-Powered Workflow Optimization for Multidisciplinary Software Development: An Automotive Industry Case Study
by: Wang, Shuai, et al.
Published: (2026)
by: Wang, Shuai, et al.
Published: (2026)
Empowering AI to Generate Better AI Code: Guided Generation of Deep Learning Projects with LLMs
by: Xie, Chen, et al.
Published: (2025)
by: Xie, Chen, et al.
Published: (2025)
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
by: Liu, Jiacheng, et al.
Published: (2026)
by: Liu, Jiacheng, et al.
Published: (2026)
PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing
by: Hou, Junyi, et al.
Published: (2025)
by: Hou, Junyi, et al.
Published: (2025)
Towards Responsible Generative AI: A Reference Architecture for Designing Foundation Model based Agents
by: Lu, Qinghua, et al.
Published: (2023)
by: Lu, Qinghua, et al.
Published: (2023)
The Code Whisperer: LLM and Graph-Based AI for Smell and Vulnerability Resolution
by: Baqar, Mohammad, et al.
Published: (2026)
by: Baqar, Mohammad, et al.
Published: (2026)
Towards Green AI: Decoding the Energy of LLM Inference in Software Development
by: Solovyeva, Lola, et al.
Published: (2026)
by: Solovyeva, Lola, et al.
Published: (2026)
ChatGPT vs. DeepSeek: A Comparative Study on AI-Based Code Generation
by: Manik, Md Motaleb Hossen
Published: (2025)
by: Manik, Md Motaleb Hossen
Published: (2025)
Similar Items
-
AutoDroid: LLM-powered Task Automation in Android
by: Wen, Hao, et al.
Published: (2023) -
Breaking Barriers in Software Testing: The Power of AI-Driven Automation
by: Naqvi, Saba, et al.
Published: (2025) -
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
by: Son, Ha Min, et al.
Published: (2025) -
On the Adoption of AI Coding Agents in Open-source Android and iOS Development
by: Khan, Muhammad Ahmad, et al.
Published: (2026) -
Understanding the Weakness of Large Language Model Agents within a Complex Android Environment
by: Xing, Mingzhe, et al.
Published: (2024)