TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Frank F., Song, Yufan, Li, Boxuan, Tang, Yuxuan, Jain, Kritanjali, Bao, Mengxue, Wang, Zora Z., Zhou, Xuhui, Guo, Zhitong, Cao, Murong, Yang, Mingyang, Lu, Hao Yang, Martin, Amaad, Su, Zhe, Maben, Leander, Mehta, Raj, Chi, Wayne, Jang, Lawrence, Xie, Yiqing, Zhou, Shuyan, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TOM-SWE: User Mental Modeling For Software Engineering Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026)
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026)
Beyond Browsing: API-Based Web Agents
von: Song, Yueqi, et al.
Veröffentlicht: (2024)
von: Song, Yueqi, et al.
Veröffentlicht: (2024)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
von: Huq, Faria, et al.
Veröffentlicht: (2025)
von: Huq, Faria, et al.
Veröffentlicht: (2025)
Agent Workflow Memory
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
Modeling Distinct Human Interaction in Web Agents
von: Huq, Faria, et al.
Veröffentlicht: (2026)
von: Huq, Faria, et al.
Veröffentlicht: (2026)
AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
von: Maben, Leander Melroy, et al.
Veröffentlicht: (2025)
von: Maben, Leander Melroy, et al.
Veröffentlicht: (2025)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
Training Proactive and Personalized LLM Agents
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
A Survey of Large Language Model Agents for Question Answering
von: Yue, Murong
Veröffentlicht: (2025)
von: Yue, Murong
Veröffentlicht: (2025)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
WebArena: A Realistic Web Environment for Building Autonomous Agents
von: Zhou, Shuyan, et al.
Veröffentlicht: (2023)
von: Zhou, Shuyan, et al.
Veröffentlicht: (2023)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2025)
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2025)
Healing the Healers: Fifty Years of Global Challenges and Progress in Nurse Psychological Wellbeing
von: Jill Maben
Veröffentlicht: (2025)
von: Jill Maben
Veröffentlicht: (2025)
Magritte [Documental] : Monsieur René Magritte / Adrian Maben, director ; Michelle Arnald y Reiner Morits, productores
von: Maben, Adrian
Veröffentlicht: (1978)
von: Maben, Adrian
Veröffentlicht: (1978)
Fittingness and Consequentialism
von: Brad Hooker
Veröffentlicht: (2026)
von: Brad Hooker
Veröffentlicht: (2026)
Double Distillation Network for Multi-Agent Reinforcement Learning
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
von: Wang, Xingyao, et al.
Veröffentlicht: (2025)
von: Wang, Xingyao, et al.
Veröffentlicht: (2025)
Coding Agents are Effective Long-Context Processors
von: Cao, Weili, et al.
Veröffentlicht: (2026)
von: Cao, Weili, et al.
Veröffentlicht: (2026)
Automated Cervical Cancer Detection through Visual Inspection with Acetic Acid in Resource-Poor Settings with Lightweight Deep Learning Models Deployed on an Android Device
von: Maben, Leander Melroy, et al.
Veröffentlicht: (2025)
von: Maben, Leander Melroy, et al.
Veröffentlicht: (2025)
How Well Does Agent Development Reflect Real-World Work?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2026)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2026)
Effective Strategies for Asynchronous Software Engineering Agents
von: Geng, Jiayi, et al.
Veröffentlicht: (2026)
von: Geng, Jiayi, et al.
Veröffentlicht: (2026)
Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale
von: Ou, Tianyue, et al.
Veröffentlicht: (2024)
von: Ou, Tianyue, et al.
Veröffentlicht: (2024)
Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
von: Zhang, Minxing, et al.
Veröffentlicht: (2025)
von: Zhang, Minxing, et al.
Veröffentlicht: (2025)
Go-Browse: Training Web Agents with Structured Exploration
von: Gandhi, Apurva, et al.
Veröffentlicht: (2025)
von: Gandhi, Apurva, et al.
Veröffentlicht: (2025)
Heterogeneous Value Decomposition Policy Fusion for Multi-Agent Cooperation
von: Wang, Siying, et al.
Veröffentlicht: (2025)
von: Wang, Siying, et al.
Veröffentlicht: (2025)
Learning Personalized Agents from Human Feedback
von: Liang, Kaiqu, et al.
Veröffentlicht: (2026)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2026)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2023)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2023)
Inducing Programmatic Skills for Agentic Tasks
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
Consequentialism, Welfarism, and Meaning in Life
von: Chad Mason Stevenson
Veröffentlicht: (2024)
von: Chad Mason Stevenson
Veröffentlicht: (2024)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
Open Government Data and Corporate Tax Avoidance: Evidence From Listed Companies of China
von: Hua Wang, et al.
Veröffentlicht: (2025)
von: Hua Wang, et al.
Veröffentlicht: (2025)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
Training Versatile Coding Agents in Synthetic Environments
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
Recent Advances in Hypergraph Neural Networks
von: Yang, Murong, et al.
Veröffentlicht: (2025)
von: Yang, Murong, et al.
Veröffentlicht: (2025)
Gym-Anything: Turn any Software into an Agent Environment
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2025)
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2025)
Chapter 4 Consequentialism and the Law in Medicine
von: Savulescu, Julian, et al.
Veröffentlicht: (2021)
von: Savulescu, Julian, et al.
Veröffentlicht: (2021)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TOM-SWE: User Mental Modeling For Software Engineering Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025) -
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025) -
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026) -
Beyond Browsing: API-Based Web Agents
von: Song, Yueqi, et al.
Veröffentlicht: (2024) -
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
von: Huq, Faria, et al.
Veröffentlicht: (2025)