AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Trivedi, Harsh, Khot, Tushar, Hartmann, Mareike, Manku, Ruskin, Dong, Vinty, Li, Edward, Gupta, Shashank, Sabharwal, Ashish, Balasubramanian, Niranjan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging In-Context Learning for Language Model Agents
von: Gupta, Shivanshu, et al.
Veröffentlicht: (2025)
von: Gupta, Shivanshu, et al.
Veröffentlicht: (2025)
ADaPT: As-Needed Decomposition and Planning with Language Models
von: Prasad, Archiki, et al.
Veröffentlicht: (2023)
von: Prasad, Archiki, et al.
Veröffentlicht: (2023)
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
von: Bogin, Ben, et al.
Veröffentlicht: (2024)
von: Bogin, Ben, et al.
Veröffentlicht: (2024)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
von: Gupta, Shashank, et al.
Veröffentlicht: (2023)
von: Gupta, Shashank, et al.
Veröffentlicht: (2023)
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
von: Wolfson, Tomer, et al.
Veröffentlicht: (2025)
von: Wolfson, Tomer, et al.
Veröffentlicht: (2025)
EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
von: Manku, Ruskin Raj, et al.
Veröffentlicht: (2025)
von: Manku, Ruskin Raj, et al.
Veröffentlicht: (2025)
Leveraging Code to Improve In-context Learning for Semantic Parsing
von: Bogin, Ben, et al.
Veröffentlicht: (2023)
von: Bogin, Ben, et al.
Veröffentlicht: (2023)
A Survey on Complex Tasks for Goal-Directed Interactive Agents
von: Hartmann, Mareike, et al.
Veröffentlicht: (2024)
von: Hartmann, Mareike, et al.
Veröffentlicht: (2024)
MMTABREAL: Real-World Benchmark for Multimodal Table Understanding
von: Titiya, Prasham, et al.
Veröffentlicht: (2025)
von: Titiya, Prasham, et al.
Veröffentlicht: (2025)
Principios de química inorg nica / G. S. Manku ; traducción de Raymundo Cea Olivares
von: Manku G. S
Veröffentlicht: (1988)
von: Manku G. S
Veröffentlicht: (1988)
Principios de química inorg nica / G. S. Maku ; traductor Raymundo Cea Olivares
von: Manku, G. S
von: Manku, G. S
MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
MapVerse: A Benchmark for Geospatial Question Answering on Diverse Real-World Maps
von: Bhat, Sharat, et al.
Veröffentlicht: (2026)
von: Bhat, Sharat, et al.
Veröffentlicht: (2026)
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
von: Kim, Joongwon, et al.
Veröffentlicht: (2024)
von: Kim, Joongwon, et al.
Veröffentlicht: (2024)
SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking
von: Liu, Guohong, et al.
Veröffentlicht: (2026)
von: Liu, Guohong, et al.
Veröffentlicht: (2026)
A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers
von: Merrill, William, et al.
Veröffentlicht: (2025)
von: Merrill, William, et al.
Veröffentlicht: (2025)
Exact Expressive Power of Transformers with Padding
von: Merrill, William, et al.
Veröffentlicht: (2025)
von: Merrill, William, et al.
Veröffentlicht: (2025)
The Expressive Power of Transformers with Chain of Thought
von: Merrill, William, et al.
Veröffentlicht: (2023)
von: Merrill, William, et al.
Veröffentlicht: (2023)
A Logic for Expressing Log-Precision Transformers
von: Merrill, William, et al.
Veröffentlicht: (2022)
von: Merrill, William, et al.
Veröffentlicht: (2022)
On the Reasoning Abilities of Masked Diffusion Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2025)
von: Svete, Anej, et al.
Veröffentlicht: (2025)
Safe Equilibrium Policy Optimization for Strategic Agent Policies
von: Arumugam, Karthika, et al.
Veröffentlicht: (2026)
von: Arumugam, Karthika, et al.
Veröffentlicht: (2026)
Certificates without Electrons? Theory and Evidence on Impacts from AI-Driven Power Demand
von: Golden, Dana, et al.
Veröffentlicht: (2026)
von: Golden, Dana, et al.
Veröffentlicht: (2026)
Efficient Generation of Diverse Cooperative Agents with World Models
von: Loo, Yi, et al.
Veröffentlicht: (2025)
von: Loo, Yi, et al.
Veröffentlicht: (2025)
ng-reactive-lint: Smarter Linting for Angular Apps
von: Balasubramanian, Shrinivass Arunachalam
Veröffentlicht: (2025)
von: Balasubramanian, Shrinivass Arunachalam
Veröffentlicht: (2025)
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences
von: Hasan, Mohammad Saqib, et al.
Veröffentlicht: (2025)
von: Hasan, Mohammad Saqib, et al.
Veröffentlicht: (2025)
Procedural Environment Generation for Tool-Use Agents
von: Sullivan, Michael, et al.
Veröffentlicht: (2025)
von: Sullivan, Michael, et al.
Veröffentlicht: (2025)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
von: Guo, JunJia, et al.
Veröffentlicht: (2026)
von: Guo, JunJia, et al.
Veröffentlicht: (2026)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
von: Shen, Chihao, et al.
Veröffentlicht: (2025)
von: Shen, Chihao, et al.
Veröffentlicht: (2025)
WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
ViMo: A Generative Visual GUI World Model for App Agents
von: Luo, Dezhao, et al.
Veröffentlicht: (2025)
von: Luo, Dezhao, et al.
Veröffentlicht: (2025)
Understanding the Logic of Direct Preference Alignment through Logic
von: Richardson, Kyle, et al.
Veröffentlicht: (2024)
von: Richardson, Kyle, et al.
Veröffentlicht: (2024)
The Illusion of State in State-Space Models
von: Merrill, William, et al.
Veröffentlicht: (2024)
von: Merrill, William, et al.
Veröffentlicht: (2024)
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
von: Chu, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Chu, Zhaoyang, et al.
Veröffentlicht: (2026)
Towards Adaptable and Interactive Image Captioning with Data Augmentation and Episodic Memory
von: Anagnostopoulou, Aliki, et al.
Veröffentlicht: (2023)
von: Anagnostopoulou, Aliki, et al.
Veröffentlicht: (2023)
Coding Agent Is Good As World Simulator
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
There's an App for That: Exploring the Market for Contraceptive Fertility Tracking Apps in the Philippines
von: Kendal Danna, et al.
Veröffentlicht: (2024)
von: Kendal Danna, et al.
Veröffentlicht: (2024)
Planet App: Kids' Book Apps Are Everywhere. But Are They Any Good?
von: Bird, Elizabeth
Veröffentlicht: (2011)
von: Bird, Elizabeth
Veröffentlicht: (2011)
The App Squad: SLJ's Advisors Weigh in on Kids' Book Apps
von: Ishizuka, Kathy
Veröffentlicht: (2011)
von: Ishizuka, Kathy
Veröffentlicht: (2011)
Ähnliche Einträge
-
Leveraging In-Context Learning for Language Model Agents
von: Gupta, Shivanshu, et al.
Veröffentlicht: (2025) -
ADaPT: As-Needed Decomposition and Planning with Language Models
von: Prasad, Archiki, et al.
Veröffentlicht: (2023) -
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
von: Bogin, Ben, et al.
Veröffentlicht: (2024) -
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
von: Gupta, Shashank, et al.
Veröffentlicht: (2023) -
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
von: Wolfson, Tomer, et al.
Veröffentlicht: (2025)