How can we assess human-agent interactions? Case studies in software agent design
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Valerie, Malhotra, Rohit, Wang, Xingyao, Michelini, Juan, Zhou, Xuhui, Soni, Aditya Bharat, Tran, Hoang H., Smith, Calvin, Talwalkar, Ameet, Neubig, Graham |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
por: Chen, Valerie, et al.
Publicado: (2025)
por: Chen, Valerie, et al.
Publicado: (2025)
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
por: Wang, Xingyao, et al.
Publicado: (2025)
por: Wang, Xingyao, et al.
Publicado: (2025)
Do LLMs exhibit human-like response biases? A case study in survey design
por: Tjuatja, Lindia, et al.
Publicado: (2023)
por: Tjuatja, Lindia, et al.
Publicado: (2023)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
por: Soni, Aditya Bharat, et al.
Publicado: (2025)
por: Soni, Aditya Bharat, et al.
Publicado: (2025)
OpenHands/software-agent-sdk: v1.21.0
por: Xingyao Wang, et al.
Publicado: (2026)
por: Xingyao Wang, et al.
Publicado: (2026)
OpenHands/software-agent-sdk: v1.19.1
por: Xingyao Wang, et al.
Publicado: (2026)
por: Xingyao Wang, et al.
Publicado: (2026)
Multitask Learning Can Improve Worst-Group Outcomes
por: Kulkarni, Atharva, et al.
Publicado: (2023)
por: Kulkarni, Atharva, et al.
Publicado: (2023)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
A Rubric-Supervised Critic from Sparse Real-World Outcomes
por: Wang, Xingyao, et al.
Publicado: (2026)
por: Wang, Xingyao, et al.
Publicado: (2026)
TOM-SWE: User Mental Modeling For Software Engineering Agents
por: Zhou, Xuhui, et al.
Publicado: (2025)
por: Zhou, Xuhui, et al.
Publicado: (2025)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
por: Kolawole, Steven, et al.
Publicado: (2024)
por: Kolawole, Steven, et al.
Publicado: (2024)
Why Do Decision Makers (Not) Use AI? A Cross-Domain Analysis of Factors Impacting AI Adoption
por: Yu, Rebecca, et al.
Publicado: (2025)
por: Yu, Rebecca, et al.
Publicado: (2025)
UPS: Efficiently Building Foundation Models for PDE Solving via Cross-Modal Adaptation
por: Shen, Junhong, et al.
Publicado: (2024)
por: Shen, Junhong, et al.
Publicado: (2024)
The Impact of Element Ordering on LM Agent Performance
por: Chi, Wayne, et al.
Publicado: (2024)
por: Chi, Wayne, et al.
Publicado: (2024)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
por: Chi, Wayne, et al.
Publicado: (2025)
por: Chi, Wayne, et al.
Publicado: (2025)
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
por: Chen, Valerie, et al.
Publicado: (2026)
por: Chen, Valerie, et al.
Publicado: (2026)
AI agents can coordinate beyond human scale
por: De Marzo, Giordano, et al.
Publicado: (2024)
por: De Marzo, Giordano, et al.
Publicado: (2024)
Need Help? Designing Proactive AI Assistants for Programming
por: Chen, Valerie, et al.
Publicado: (2024)
por: Chen, Valerie, et al.
Publicado: (2024)
CodingGenie: A Proactive LLM-Powered Programming Assistant
por: Zhao, Sebastian, et al.
Publicado: (2025)
por: Zhao, Sebastian, et al.
Publicado: (2025)
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
por: Pan, Jane, et al.
Publicado: (2025)
por: Pan, Jane, et al.
Publicado: (2025)
Agreement-Based Cascading for Efficient Inference
por: Kolawole, Steven, et al.
Publicado: (2024)
por: Kolawole, Steven, et al.
Publicado: (2024)
Training Proactive and Personalized LLM Agents
por: Sun, Weiwei, et al.
Publicado: (2025)
por: Sun, Weiwei, et al.
Publicado: (2025)
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
por: Soni, Aditya Bharat, et al.
Publicado: (2026)
por: Soni, Aditya Bharat, et al.
Publicado: (2026)
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
por: Sutawika, Lintang, et al.
Publicado: (2026)
por: Sutawika, Lintang, et al.
Publicado: (2026)
Provably tuning the ElasticNet across instances
por: Balcan, Maria-Florina, et al.
Publicado: (2022)
por: Balcan, Maria-Florina, et al.
Publicado: (2022)
Learning to Relax: Setting Solver Parameters Across a Sequence of Linear System Instances
por: Khodak, Mikhail, et al.
Publicado: (2023)
por: Khodak, Mikhail, et al.
Publicado: (2023)
Where Does My Model Underperform? A Human Evaluation of Slice Discovery Algorithms
por: Johnson, Nari, et al.
Publicado: (2023)
por: Johnson, Nari, et al.
Publicado: (2023)
Comparing Developer and LLM Biases in Code Evaluation
por: Mittal, Aditya, et al.
Publicado: (2026)
por: Mittal, Aditya, et al.
Publicado: (2026)
How can AI agents support journalists' work? An experiment with designing an LLM-driven intelligent reporting system
por: Maltezos, Vasileios, et al.
Publicado: (2025)
por: Maltezos, Vasileios, et al.
Publicado: (2025)
CoMind: Towards Community-Driven Agents for Machine Learning Engineering
por: Li, Sijie, et al.
Publicado: (2025)
por: Li, Sijie, et al.
Publicado: (2025)
FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
por: Feng, Shengyu, et al.
Publicado: (2025)
por: Feng, Shengyu, et al.
Publicado: (2025)
Learning Personalized Decision Support Policies
por: Bhatt, Umang, et al.
Publicado: (2023)
por: Bhatt, Umang, et al.
Publicado: (2023)
What we owe to impaired agents
por: Giacomo Floris
Publicado: (2024)
por: Giacomo Floris
Publicado: (2024)
How Gaming Could Improve Information Literacy
por: Doshi, Ameet
Publicado: (2006)
por: Doshi, Ameet
Publicado: (2006)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
por: Song, Yueqi, et al.
Publicado: (2024)
por: Song, Yueqi, et al.
Publicado: (2024)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevance
por: Khanuja, Simran, et al.
Publicado: (2024)
por: Khanuja, Simran, et al.
Publicado: (2024)
Asymmetric interaction preference induces cooperation in human-agent hybrid game
por: Jia, Danyang, et al.
Publicado: (2024)
por: Jia, Danyang, et al.
Publicado: (2024)
Evaluando características del agente software
por: Héctor Soza Pollman
Publicado: (2014)
por: Héctor Soza Pollman
Publicado: (2014)
How inventory consignment programs can improve supply chain performance: a process oriented perspective
por: Manoj K. Malhotra
Publicado: (2017)
por: Manoj K. Malhotra
Publicado: (2017)
Ejemplares similares
-
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
por: Chen, Valerie, et al.
Publicado: (2025) -
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
por: Wang, Xingyao, et al.
Publicado: (2025) -
Do LLMs exhibit human-like response biases? A case study in survey design
por: Tjuatja, Lindia, et al.
Publicado: (2023) -
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
por: Soni, Aditya Bharat, et al.
Publicado: (2025) -
OpenHands/software-agent-sdk: v1.21.0
por: Xingyao Wang, et al.
Publicado: (2026)