How Well Does Agent Development Reflect Real-World Work?
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Zora Zhiruo, Vijayvargiya, Sanidhya, Chen, Aspen, Zhang, Hanmo, Arangarajan, Venu Arvind, Chen, Jett, Chen, Valerie, Yang, Diyi, Fried, Daniel, Neubig, Graham |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
por: Wang, Zora Zhiruo, et al.
Publicado: (2025)
por: Wang, Zora Zhiruo, et al.
Publicado: (2025)
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2026)
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2026)
Agent Workflow Memory
por: Wang, Zora Zhiruo, et al.
Publicado: (2024)
por: Wang, Zora Zhiruo, et al.
Publicado: (2024)
Modeling Distinct Human Interaction in Web Agents
por: Huq, Faria, et al.
Publicado: (2026)
por: Huq, Faria, et al.
Publicado: (2026)
Inducing Programmatic Skills for Agentic Tasks
por: Wang, Zora Zhiruo, et al.
Publicado: (2025)
por: Wang, Zora Zhiruo, et al.
Publicado: (2025)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
Efficient On-Device Agents via Adaptive Context Management
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
TOM-SWE: User Mental Modeling For Software Engineering Agents
por: Zhou, Xuhui, et al.
Publicado: (2025)
por: Zhou, Xuhui, et al.
Publicado: (2025)
TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
por: Wang, Zhiruo, et al.
Publicado: (2024)
por: Wang, Zhiruo, et al.
Publicado: (2024)
Benchmarking Failures in Tool-Augmented Language Models
por: Treviño, Eduardo, et al.
Publicado: (2025)
por: Treviño, Eduardo, et al.
Publicado: (2025)
What Are Tools Anyway? A Survey from the Language Model Perspective
por: Wang, Zhiruo, et al.
Publicado: (2024)
por: Wang, Zhiruo, et al.
Publicado: (2024)
A Rubric-Supervised Critic from Sparse Real-World Outcomes
por: Wang, Xingyao, et al.
Publicado: (2026)
por: Wang, Xingyao, et al.
Publicado: (2026)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
por: Wang, Zora Zhiruo, et al.
Publicado: (2024)
por: Wang, Zora Zhiruo, et al.
Publicado: (2024)
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
por: Chen, Valerie, et al.
Publicado: (2025)
por: Chen, Valerie, et al.
Publicado: (2025)
ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
por: Waghjale, Siddhant, et al.
Publicado: (2024)
por: Waghjale, Siddhant, et al.
Publicado: (2024)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
por: Huq, Faria, et al.
Publicado: (2025)
por: Huq, Faria, et al.
Publicado: (2025)
RAGGED: Towards Informed Design of Scalable and Stable RAG Systems
por: Hsia, Jennifer, et al.
Publicado: (2024)
por: Hsia, Jennifer, et al.
Publicado: (2024)
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
por: Sutawika, Lintang, et al.
Publicado: (2026)
por: Sutawika, Lintang, et al.
Publicado: (2026)
AutoPresent: Designing Structured Visuals from Scratch
por: Ge, Jiaxin, et al.
Publicado: (2025)
por: Ge, Jiaxin, et al.
Publicado: (2025)
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
por: Padmakumar, Vishakh, et al.
Publicado: (2026)
por: Padmakumar, Vishakh, et al.
Publicado: (2026)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
por: Soni, Aditya Bharat, et al.
Publicado: (2025)
por: Soni, Aditya Bharat, et al.
Publicado: (2025)
SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
por: Zheng, Boyuan, et al.
Publicado: (2025)
por: Zheng, Boyuan, et al.
Publicado: (2025)
ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
por: Xiao, Yunzhong, et al.
Publicado: (2025)
por: Xiao, Yunzhong, et al.
Publicado: (2025)
API-Assisted Code Generation for Question Answering on Varied Table Structures
por: Cao, Yihan, et al.
Publicado: (2023)
por: Cao, Yihan, et al.
Publicado: (2023)
Effective Strategies for Asynchronous Software Engineering Agents
por: Geng, Jiayi, et al.
Publicado: (2026)
por: Geng, Jiayi, et al.
Publicado: (2026)
Go-Browse: Training Web Agents with Structured Exploration
por: Gandhi, Apurva, et al.
Publicado: (2025)
por: Gandhi, Apurva, et al.
Publicado: (2025)
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
por: Xu, Frank F., et al.
Publicado: (2024)
por: Xu, Frank F., et al.
Publicado: (2024)
Do LLMs exhibit human-like response biases? A case study in survey design
por: Tjuatja, Lindia, et al.
Publicado: (2023)
por: Tjuatja, Lindia, et al.
Publicado: (2023)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
por: Song, Yueqi, et al.
Publicado: (2024)
por: Song, Yueqi, et al.
Publicado: (2024)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
por: Chi, Wayne, et al.
Publicado: (2025)
por: Chi, Wayne, et al.
Publicado: (2025)
Authorship and Text-making in Early China
por: Zhang, Hanmo
Publicado: (2022)
por: Zhang, Hanmo
Publicado: (2022)
Training Versatile Coding Agents in Synthetic Environments
por: Zhu, Yiqi, et al.
Publicado: (2025)
por: Zhu, Yiqi, et al.
Publicado: (2025)
Repetition Improves Language Model Embeddings
por: Springer, Jacob Mitchell, et al.
Publicado: (2024)
por: Springer, Jacob Mitchell, et al.
Publicado: (2024)
Gym-Anything: Turn any Software into an Agent Environment
por: Aggarwal, Pranjal, et al.
Publicado: (2026)
por: Aggarwal, Pranjal, et al.
Publicado: (2026)
How Well Does Readily Available Literature Reflect the Changing American Family?
por: Obert, Beth
Publicado: (1988)
por: Obert, Beth
Publicado: (1988)
DECEPTICON: How Dark Patterns Manipulate Web Agents
por: Cuvin, Phil, et al.
Publicado: (2025)
por: Cuvin, Phil, et al.
Publicado: (2025)
Task Shifting, eHealth and Shared Decision‐Making—Preference Heterogeneity in the Adult Population for Developments in Outpatient Primary Healthcare
por: Zora Föhn
Publicado: (2025)
por: Zora Föhn
Publicado: (2025)
Expanding conceptualizations of engineering persistence: Examining four undergraduate Black men's dual‐degree experiences
por: Christopher C. Jett
Publicado: (2025)
por: Christopher C. Jett
Publicado: (2025)
Enhancing Workplace Productivity and Well-being Using AI Agent
por: K, Ravirajan, et al.
Publicado: (2025)
por: K, Ravirajan, et al.
Publicado: (2025)
Ejemplares similares
-
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025) -
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
por: Wang, Zora Zhiruo, et al.
Publicado: (2025) -
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2026) -
Agent Workflow Memory
por: Wang, Zora Zhiruo, et al.
Publicado: (2024) -
Modeling Distinct Human Interaction in Web Agents
por: Huq, Faria, et al.
Publicado: (2026)