How Well Does Agent Development Reflect Real-World Work?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zora Zhiruo, Vijayvargiya, Sanidhya, Chen, Aspen, Zhang, Hanmo, Arangarajan, Venu Arvind, Chen, Jett, Chen, Valerie, Yang, Diyi, Fried, Daniel, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2026)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2026)
Agent Workflow Memory
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
Modeling Distinct Human Interaction in Web Agents
von: Huq, Faria, et al.
Veröffentlicht: (2026)
von: Huq, Faria, et al.
Veröffentlicht: (2026)
Inducing Programmatic Skills for Agentic Tasks
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
Efficient On-Device Agents via Adaptive Context Management
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
TOM-SWE: User Mental Modeling For Software Engineering Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
von: Wang, Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zhiruo, et al.
Veröffentlicht: (2024)
Benchmarking Failures in Tool-Augmented Language Models
von: Treviño, Eduardo, et al.
Veröffentlicht: (2025)
von: Treviño, Eduardo, et al.
Veröffentlicht: (2025)
What Are Tools Anyway? A Survey from the Language Model Perspective
von: Wang, Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zhiruo, et al.
Veröffentlicht: (2024)
A Rubric-Supervised Critic from Sparse Real-World Outcomes
von: Wang, Xingyao, et al.
Veröffentlicht: (2026)
von: Wang, Xingyao, et al.
Veröffentlicht: (2026)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
von: Chen, Valerie, et al.
Veröffentlicht: (2025)
von: Chen, Valerie, et al.
Veröffentlicht: (2025)
ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
von: Waghjale, Siddhant, et al.
Veröffentlicht: (2024)
von: Waghjale, Siddhant, et al.
Veröffentlicht: (2024)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
von: Huq, Faria, et al.
Veröffentlicht: (2025)
von: Huq, Faria, et al.
Veröffentlicht: (2025)
RAGGED: Towards Informed Design of Scalable and Stable RAG Systems
von: Hsia, Jennifer, et al.
Veröffentlicht: (2024)
von: Hsia, Jennifer, et al.
Veröffentlicht: (2024)
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026)
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026)
AutoPresent: Designing Structured Visuals from Scratch
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2026)
von: Padmakumar, Vishakh, et al.
Veröffentlicht: (2026)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2025)
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2025)
SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
von: Xiao, Yunzhong, et al.
Veröffentlicht: (2025)
von: Xiao, Yunzhong, et al.
Veröffentlicht: (2025)
API-Assisted Code Generation for Question Answering on Varied Table Structures
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
Effective Strategies for Asynchronous Software Engineering Agents
von: Geng, Jiayi, et al.
Veröffentlicht: (2026)
von: Geng, Jiayi, et al.
Veröffentlicht: (2026)
Go-Browse: Training Web Agents with Structured Exploration
von: Gandhi, Apurva, et al.
Veröffentlicht: (2025)
von: Gandhi, Apurva, et al.
Veröffentlicht: (2025)
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
von: Xu, Frank F., et al.
Veröffentlicht: (2024)
von: Xu, Frank F., et al.
Veröffentlicht: (2024)
Do LLMs exhibit human-like response biases? A case study in survey design
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2023)
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2023)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
von: Song, Yueqi, et al.
Veröffentlicht: (2024)
von: Song, Yueqi, et al.
Veröffentlicht: (2024)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
von: Chi, Wayne, et al.
Veröffentlicht: (2025)
von: Chi, Wayne, et al.
Veröffentlicht: (2025)
Authorship and Text-making in Early China
von: Zhang, Hanmo
Veröffentlicht: (2022)
von: Zhang, Hanmo
Veröffentlicht: (2022)
Training Versatile Coding Agents in Synthetic Environments
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
Repetition Improves Language Model Embeddings
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2024)
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2024)
Gym-Anything: Turn any Software into an Agent Environment
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
How Well Does Readily Available Literature Reflect the Changing American Family?
von: Obert, Beth
Veröffentlicht: (1988)
von: Obert, Beth
Veröffentlicht: (1988)
DECEPTICON: How Dark Patterns Manipulate Web Agents
von: Cuvin, Phil, et al.
Veröffentlicht: (2025)
von: Cuvin, Phil, et al.
Veröffentlicht: (2025)
Task Shifting, eHealth and Shared Decision‐Making—Preference Heterogeneity in the Adult Population for Developments in Outpatient Primary Healthcare
von: Zora Föhn
Veröffentlicht: (2025)
von: Zora Föhn
Veröffentlicht: (2025)
Expanding conceptualizations of engineering persistence: Examining four undergraduate Black men's dual‐degree experiences
von: Christopher C. Jett
Veröffentlicht: (2025)
von: Christopher C. Jett
Veröffentlicht: (2025)
Enhancing Workplace Productivity and Well-being Using AI Agent
von: K, Ravirajan, et al.
Veröffentlicht: (2025)
von: K, Ravirajan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025) -
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025) -
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2026) -
Agent Workflow Memory
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024) -
Modeling Distinct Human Interaction in Web Agents
von: Huq, Faria, et al.
Veröffentlicht: (2026)