Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
Fuente:
arXiv
Salvato in:
| Autori principali: | Vijayvargiya, Sanidhya, Zhou, Xuhui, Yerukola, Akhila, Sap, Maarten, Neubig, Graham |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2026)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2026)
TOM-SWE: User Mental Modeling For Software Engineering Agents
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
Social World Models
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
di: Yerukola, Akhila, et al.
Pubblicazione: (2024)
di: Yerukola, Akhila, et al.
Pubblicazione: (2024)
Efficient On-Device Agents via Adaptive Context Management
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
di: Yerukola, Akhila, et al.
Pubblicazione: (2025)
di: Yerukola, Akhila, et al.
Pubblicazione: (2025)
Effective Strategies for Asynchronous Software Engineering Agents
di: Geng, Jiayi, et al.
Pubblicazione: (2026)
di: Geng, Jiayi, et al.
Pubblicazione: (2026)
Training Proactive and Personalized LLM Agents
di: Sun, Weiwei, et al.
Pubblicazione: (2025)
di: Sun, Weiwei, et al.
Pubblicazione: (2025)
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)
How Well Does Agent Development Reflect Real-World Work?
di: Wang, Zora Zhiruo, et al.
Pubblicazione: (2026)
di: Wang, Zora Zhiruo, et al.
Pubblicazione: (2026)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
di: Zhou, Xuhui, et al.
Pubblicazione: (2023)
di: Zhou, Xuhui, et al.
Pubblicazione: (2023)
1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
di: Li, Wenkai, et al.
Pubblicazione: (2025)
di: Li, Wenkai, et al.
Pubblicazione: (2025)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
di: Zhou, Xuhui, et al.
Pubblicazione: (2024)
di: Zhou, Xuhui, et al.
Pubblicazione: (2024)
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
di: Shen, Jocelyn, et al.
Pubblicazione: (2025)
di: Shen, Jocelyn, et al.
Pubblicazione: (2025)
Gym-Anything: Turn any Software into an Agent Environment
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2026)
di: Aggarwal, Pranjal, et al.
Pubblicazione: (2026)
SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling
di: Wang, Haoran, et al.
Pubblicazione: (2025)
di: Wang, Haoran, et al.
Pubblicazione: (2025)
SWE-smith: Scaling Data for Software Engineering Agents
di: Yang, John, et al.
Pubblicazione: (2025)
di: Yang, John, et al.
Pubblicazione: (2025)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
di: Mun, Jimin, et al.
Pubblicazione: (2026)
di: Mun, Jimin, et al.
Pubblicazione: (2026)
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
di: Zhou, Xuhui, et al.
Pubblicazione: (2026)
di: Zhou, Xuhui, et al.
Pubblicazione: (2026)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
di: Liang, Jiarong, et al.
Pubblicazione: (2026)
di: Liang, Jiarong, et al.
Pubblicazione: (2026)
Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
di: Wang, Qiaosi, et al.
Pubblicazione: (2025)
di: Wang, Qiaosi, et al.
Pubblicazione: (2025)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
di: Su, Zhe, et al.
Pubblicazione: (2024)
di: Su, Zhe, et al.
Pubblicazione: (2024)
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
di: Wang, Xingyao, et al.
Pubblicazione: (2025)
di: Wang, Xingyao, et al.
Pubblicazione: (2025)
On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
di: Huang, Jen-tse, et al.
Pubblicazione: (2024)
di: Huang, Jen-tse, et al.
Pubblicazione: (2024)
Out of Style: RAG's Fragility to Linguistic Variation
di: Cao, Tianyu, et al.
Pubblicazione: (2025)
di: Cao, Tianyu, et al.
Pubblicazione: (2025)
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
di: Rao, Abhinav, et al.
Pubblicazione: (2024)
di: Rao, Abhinav, et al.
Pubblicazione: (2024)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
di: Ding, Yifeng, et al.
Pubblicazione: (2026)
di: Ding, Yifeng, et al.
Pubblicazione: (2026)
SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
di: Fan, Xianzhe, et al.
Pubblicazione: (2025)
di: Fan, Xianzhe, et al.
Pubblicazione: (2025)
Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents
di: Stoisser, Josefa Lia, et al.
Pubblicazione: (2026)
di: Stoisser, Josefa Lia, et al.
Pubblicazione: (2026)
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
di: Zeng, Liang, et al.
Pubblicazione: (2025)
di: Zeng, Liang, et al.
Pubblicazione: (2025)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
di: Han, Tingxu, et al.
Pubblicazione: (2026)
di: Han, Tingxu, et al.
Pubblicazione: (2026)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
di: Yuan, Danlong, et al.
Pubblicazione: (2026)
di: Yuan, Danlong, et al.
Pubblicazione: (2026)
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
di: Xia, Chunqiu Steven, et al.
Pubblicazione: (2025)
di: Xia, Chunqiu Steven, et al.
Pubblicazione: (2025)
Training Software Engineering Agents and Verifiers with SWE-Gym
di: Pan, Jiayi, et al.
Pubblicazione: (2024)
di: Pan, Jiayi, et al.
Pubblicazione: (2024)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
di: Yang, John, et al.
Pubblicazione: (2024)
di: Yang, John, et al.
Pubblicazione: (2024)
Ambig-IaC: Multi-level Disambiguation for Interactive Cloud Infrastructure-as-Code Synthesis
di: Yang, Zhenning, et al.
Pubblicazione: (2026)
di: Yang, Zhenning, et al.
Pubblicazione: (2026)
SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks
di: Zeng, Yucheng, et al.
Pubblicazione: (2026)
di: Zeng, Yucheng, et al.
Pubblicazione: (2026)
SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
di: Antoniades, Antonis, et al.
Pubblicazione: (2024)
di: Antoniades, Antonis, et al.
Pubblicazione: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2026) -
TOM-SWE: User Mental Modeling For Software Engineering Agents
di: Zhou, Xuhui, et al.
Pubblicazione: (2025) -
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2025) -
Social World Models
di: Zhou, Xuhui, et al.
Pubblicazione: (2025) -
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
di: Yerukola, Akhila, et al.
Pubblicazione: (2024)