Measuring Agents in Production
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915771097022464 |
|---|---|
| author | Pan, Melissa Z. Arabzadeh, Negar Cogo, Riccardo Zhu, Yuxuan Xiong, Alexander Agrawal, Lakshya A Mao, Huanzhi Shen, Emma Pallerla, Sid Patel, Liana Liu, Shu Shi, Tianneng Liu, Xiaoyuan Davis, Jared Quincy Lacavalla, Emmanuele Basile, Alessandro Yang, Shuyi Castro, Paul Kang, Daniel Gonzalez, Joseph E. Sen, Koushik Song, Dawn Stoica, Ion Zaharia, Matei Ellis, Marquita |
| author_facet | Pan, Melissa Z. Arabzadeh, Negar Cogo, Riccardo Zhu, Yuxuan Xiong, Alexander Agrawal, Lakshya A Mao, Huanzhi Shen, Emma Pallerla, Sid Patel, Liana Liu, Shu Shi, Tianneng Liu, Xiaoyuan Davis, Jared Quincy Lacavalla, Emmanuele Basile, Alessandro Yang, Shuyi Castro, Paul Kang, Daniel Gonzalez, Joseph E. Sen, Koushik Song, Dawn Stoica, Ion Zaharia, Matei Ellis, Marquita |
| contents | LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful. We present the first systematic study of Measuring Agents in Production, MAP, using first-hand data from agent developers. We conducted 20 case studies via in-depth interviews and surveyed 306 practitioners across 26 domains. We investigate why organizations build agents, how they build them, how they evaluate them, and their top development challenges. Our study finds that production agents are built using simple, controllable approaches: 68% execute at most 10 steps before human intervention, 70% rely on prompting off-the-shelf models instead of weight tuning, and 74% depend primarily on human evaluation. Reliability (consistent correct behavior over time) remains the top development challenge, which practitioners currently address through systems-level design. MAP documents the current state of production agents, providing the research community with visibility into deployment realities and under-explored research avenues. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_04123 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Measuring Agents in Production Pan, Melissa Z. Arabzadeh, Negar Cogo, Riccardo Zhu, Yuxuan Xiong, Alexander Agrawal, Lakshya A Mao, Huanzhi Shen, Emma Pallerla, Sid Patel, Liana Liu, Shu Shi, Tianneng Liu, Xiaoyuan Davis, Jared Quincy Lacavalla, Emmanuele Basile, Alessandro Yang, Shuyi Castro, Paul Kang, Daniel Gonzalez, Joseph E. Sen, Koushik Song, Dawn Stoica, Ion Zaharia, Matei Ellis, Marquita Computers and Society Artificial Intelligence Machine Learning Software Engineering LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful. We present the first systematic study of Measuring Agents in Production, MAP, using first-hand data from agent developers. We conducted 20 case studies via in-depth interviews and surveyed 306 practitioners across 26 domains. We investigate why organizations build agents, how they build them, how they evaluate them, and their top development challenges. Our study finds that production agents are built using simple, controllable approaches: 68% execute at most 10 steps before human intervention, 70% rely on prompting off-the-shelf models instead of weight tuning, and 74% depend primarily on human evaluation. Reliability (consistent correct behavior over time) remains the top development challenge, which practitioners currently address through systems-level design. MAP documents the current state of production agents, providing the research community with visibility into deployment realities and under-explored research avenues. |
| title | Measuring Agents in Production |
| topic | Computers and Society Artificial Intelligence Machine Learning Software Engineering |
| url | https://arxiv.org/abs/2512.04123 |