Understanding the Weakness of Large Language Model Agents within a Complex Android Environment
Fuente:
arXiv
Saved in:
| Main Authors: | Xing, Mingzhe, Zhang, Rongkai, Xue, Hui, Chen, Qi, Yang, Fan, Xiao, Zhen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPROUT: an Interactive Authoring Tool for Generating Programming Tutorials with the Visualization of Large Language Models
by: Liu, Yihan, et al.
Published: (2023)
by: Liu, Yihan, et al.
Published: (2023)
Large Language Models as Visualization Agents for Immersive Binary Reverse Engineering
by: Brown, Dennis, et al.
Published: (2025)
by: Brown, Dennis, et al.
Published: (2025)
PriviSense: A Frida-Based Framework for Multi-Sensor Spoofing on Android
by: Khalilov, Ibrahim, et al.
Published: (2026)
by: Khalilov, Ibrahim, et al.
Published: (2026)
Recommending Usability Improvements with Multimodal Large Language Models
by: Lubos, Sebastian, et al.
Published: (2026)
by: Lubos, Sebastian, et al.
Published: (2026)
How Scientists Use Large Language Models to Program
by: O'Brien, Gabrielle
Published: (2025)
by: O'Brien, Gabrielle
Published: (2025)
Story Arena: A Multi-Agent Environment for Envisioning the Future of Software Engineering
by: Weisz, Justin D., et al.
Published: (2025)
by: Weisz, Justin D., et al.
Published: (2025)
Usability Analysis of Configurator User Interfaces with Multimodal Large Language Models
by: Lubos, Sebastian, et al.
Published: (2026)
by: Lubos, Sebastian, et al.
Published: (2026)
Generating Complex Code Analyzers from Natural Language Questions
by: Nazari, Amirmohammad, et al.
Published: (2026)
by: Nazari, Amirmohammad, et al.
Published: (2026)
A Model for Understanding and Reducing Developer Burnout
by: Trinkenreich, Bianca, et al.
Published: (2023)
by: Trinkenreich, Bianca, et al.
Published: (2023)
In-IDE Human-AI Experience in the Era of Large Language Models; A Literature Review
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Inferring Alt-text For UI Icons With Large Language Models During App Development
by: Haque, Sabrina, et al.
Published: (2024)
by: Haque, Sabrina, et al.
Published: (2024)
M2AR: A Web-based Modeling Environment for the Augmented Reality Workflow Modeling Language
by: Muff, Fabian, et al.
Published: (2024)
by: Muff, Fabian, et al.
Published: (2024)
Bridging the Interpretation Gap in Accessibility Testing: Empathetic and Legal-Aware Bug Report Generation via Large Language Models
by: Koyama, Ryoya, et al.
Published: (2026)
by: Koyama, Ryoya, et al.
Published: (2026)
An Exploratory Study on Upper-Level Computing Students' Use of Large Language Models as Tools in a Semester-Long Project
by: Tanay, Ben Arie, et al.
Published: (2024)
by: Tanay, Ben Arie, et al.
Published: (2024)
Exit the Code: A Model for Understanding Career Abandonment Intention Among Software Developers
by: Massoni, Tiago, et al.
Published: (2025)
by: Massoni, Tiago, et al.
Published: (2025)
Evaluating the Quality of Code Comments Generated by Large Language Models for Novice Programmers
by: Fan, Aysa Xuemo, et al.
Published: (2024)
by: Fan, Aysa Xuemo, et al.
Published: (2024)
Understanding User Mental Models in AI-Driven Code Completion Tools: Insights from an Elicitation Study
by: Desolda, Giuseppe, et al.
Published: (2025)
by: Desolda, Giuseppe, et al.
Published: (2025)
"Always Nice and Confident, Sometimes Wrong": Developer's Experiences Engaging Large Language Models (LLMs) Versus Human-Powered Q&A Platforms for Coding Support
by: Li, Jiachen, et al.
Published: (2023)
by: Li, Jiachen, et al.
Published: (2023)
Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild
by: Kumar, Aayush, et al.
Published: (2025)
by: Kumar, Aayush, et al.
Published: (2025)
SmartEx: A Framework for Generating User-Centric Explanations in Smart Environments
by: Sadeghi, Mersedeh, et al.
Published: (2024)
by: Sadeghi, Mersedeh, et al.
Published: (2024)
OSSDoorway: A Gamified Environment to Scaffold Student Contributions to Open Source Software
by: Santos, Italo, et al.
Published: (2025)
by: Santos, Italo, et al.
Published: (2025)
Towards an Understanding of Developer Experience-Driven Transparency in Software Ecosystems
by: Zacarias, Rodrigo Oliveira, et al.
Published: (2025)
by: Zacarias, Rodrigo Oliveira, et al.
Published: (2025)
HookLens: Visual Analytics for Understanding React Hooks Structures
by: Hwang, Suyeon, et al.
Published: (2026)
by: Hwang, Suyeon, et al.
Published: (2026)
ChatGPT on the Road: Leveraging Large Language Model-Powered In-vehicle Conversational Agents for Safer and More Enjoyable Driving Experience
by: Bond, Yeana Lee, et al.
Published: (2025)
by: Bond, Yeana Lee, et al.
Published: (2025)
On the Utility of Domain Modeling Assistance with Large Language Models
by: Chaaben, Meriem Ben, et al.
Published: (2024)
by: Chaaben, Meriem Ben, et al.
Published: (2024)
The Evolution of Information Seeking in Software Development: Understanding the Role and Impact of AI Assistants
by: Haque, Ebtesam Al, et al.
Published: (2024)
by: Haque, Ebtesam Al, et al.
Published: (2024)
Understanding the Human-LLM Dynamic: A Literature Survey of LLM Use in Programming Tasks
by: Etsenake, Deborah, et al.
Published: (2024)
by: Etsenake, Deborah, et al.
Published: (2024)
Understanding Documentation Use Through Log Analysis: An Exploratory Case Study of Four Cloud Services
by: Nam, Daye, et al.
Published: (2023)
by: Nam, Daye, et al.
Published: (2023)
What Pulls the Strings? Understanding the Characteristics and Role of Argumentation in Open-Source Software Usability Discussions
by: Sanei, Arghavan, et al.
Published: (2025)
by: Sanei, Arghavan, et al.
Published: (2025)
A Low-Code Approach for the Automatic Personalization of Conversational Agents
by: Conrardy, Aaron, et al.
Published: (2026)
by: Conrardy, Aaron, et al.
Published: (2026)
The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
by: Mozannar, Hussein, et al.
Published: (2024)
by: Mozannar, Hussein, et al.
Published: (2024)
"I Would Have Written My Code Differently'': Beginners Struggle to Understand LLM-Generated Code
by: Zi, Yangtian, et al.
Published: (2025)
by: Zi, Yangtian, et al.
Published: (2025)
Learning Programming in Informal Spaces: Using Emotion as a Lens to Understand Novice Struggles on r/learnprogramming
by: Hasan, Alif Al, et al.
Published: (2025)
by: Hasan, Alif Al, et al.
Published: (2025)
Investigating Conversational Agents to Support Secondary School Students Learning CSP
by: Frazier, Matthew, et al.
Published: (2026)
by: Frazier, Matthew, et al.
Published: (2026)
Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?
by: Kou, Bonan, et al.
Published: (2023)
by: Kou, Bonan, et al.
Published: (2023)
Investigating Multimodal Large Language Models to Support Usability Evaluation
by: Lubos, Sebastian, et al.
Published: (2025)
by: Lubos, Sebastian, et al.
Published: (2025)
How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging
by: Ma, Qianou, et al.
Published: (2023)
by: Ma, Qianou, et al.
Published: (2023)
AutoGraph: A Knowledge-Graph Framework for Modeling Interface Interaction and Automating Procedure Execution in Digital Nuclear Control Rooms
by: Xiao, Xingyu, et al.
Published: (2025)
by: Xiao, Xingyu, et al.
Published: (2025)
Programming by Chat: A Large-Scale Behavioral Analysis of 11,579 Real-World AI-Assisted IDE Sessions
by: Tang, Ningzhi, et al.
Published: (2026)
by: Tang, Ningzhi, et al.
Published: (2026)
Quantifying Interface Procedure Coupling Risks in Digital Nuclear Control Rooms: An Event Based Human Reliability Assessment
by: Xiao, Xingyu, et al.
Published: (2026)
by: Xiao, Xingyu, et al.
Published: (2026)
Similar Items
-
SPROUT: an Interactive Authoring Tool for Generating Programming Tutorials with the Visualization of Large Language Models
by: Liu, Yihan, et al.
Published: (2023) -
Large Language Models as Visualization Agents for Immersive Binary Reverse Engineering
by: Brown, Dennis, et al.
Published: (2025) -
PriviSense: A Frida-Based Framework for Multi-Sensor Spoofing on Android
by: Khalilov, Ibrahim, et al.
Published: (2026) -
Recommending Usability Improvements with Multimodal Large Language Models
by: Lubos, Sebastian, et al.
Published: (2026) -
How Scientists Use Large Language Models to Program
by: O'Brien, Gabrielle
Published: (2025)