OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vijayvargiya, Sanidhya, Soni, Aditya Bharat, Zhou, Xuhui, Wang, Zora Zhiruo, Dziri, Nouha, Neubig, Graham, Sap, Maarten |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
TOM-SWE: User Mental Modeling For Software Engineering Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
How Well Does Agent Development Reflect Real-World Work?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2026)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2026)
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2026)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2026)
Efficient On-Device Agents via Adaptive Context Management
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026)
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026)
Agent Workflow Memory
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2025)
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2025)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
von: Li, Jing-Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2024)
Training Proactive and Personalized LLM Agents
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2024)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2024)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
von: Huq, Faria, et al.
Veröffentlicht: (2025)
von: Huq, Faria, et al.
Veröffentlicht: (2025)
Inducing Programmatic Skills for Agentic Tasks
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2025)
AutoPresent: Designing Structured Visuals from Scratch
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2023)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2023)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
von: Bennion, Jonathan, et al.
Veröffentlicht: (2025)
von: Bennion, Jonathan, et al.
Veröffentlicht: (2025)
Benchmarking Failures in Tool-Augmented Language Models
von: Treviño, Eduardo, et al.
Veröffentlicht: (2025)
von: Treviño, Eduardo, et al.
Veröffentlicht: (2025)
Modeling Distinct Human Interaction in Web Agents
von: Huq, Faria, et al.
Veröffentlicht: (2026)
von: Huq, Faria, et al.
Veröffentlicht: (2026)
1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
von: Wang, Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zhiruo, et al.
Veröffentlicht: (2024)
SOTOPIA-TOM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind
von: YS, Yashwanth, et al.
Veröffentlicht: (2026)
von: YS, Yashwanth, et al.
Veröffentlicht: (2026)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
von: Su, Zhe, et al.
Veröffentlicht: (2024)
von: Su, Zhe, et al.
Veröffentlicht: (2024)
Social World Models
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
von: Jiang, Liwei, et al.
Veröffentlicht: (2025)
von: Jiang, Liwei, et al.
Veröffentlicht: (2025)
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
von: Xu, Frank F., et al.
Veröffentlicht: (2024)
von: Xu, Frank F., et al.
Veröffentlicht: (2024)
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2026)
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2026)
SOTOPIA-$π$: Interactive Learning of Socially Intelligent Language Agents
von: Wang, Ruiyi, et al.
Veröffentlicht: (2024)
von: Wang, Ruiyi, et al.
Veröffentlicht: (2024)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
von: Han, Seungju, et al.
Veröffentlicht: (2024)
von: Han, Seungju, et al.
Veröffentlicht: (2024)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
von: Wang, Zijun, et al.
Veröffentlicht: (2026)
von: Wang, Zijun, et al.
Veröffentlicht: (2026)
Stereotype or Personalization? User Identity Biases Chatbot Recommendations
von: Kantharuban, Anjali, et al.
Veröffentlicht: (2024)
von: Kantharuban, Anjali, et al.
Veröffentlicht: (2024)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety
von: François, Camille, et al.
Veröffentlicht: (2025)
von: François, Camille, et al.
Veröffentlicht: (2025)
ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
von: Xiao, Yunzhong, et al.
Veröffentlicht: (2025)
von: Xiao, Yunzhong, et al.
Veröffentlicht: (2025)
SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
When Should AI Read the Room? Public Perceptions of Social Intelligence in AI Agents
von: Mathur, Leena, et al.
Veröffentlicht: (2026)
von: Mathur, Leena, et al.
Veröffentlicht: (2026)
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
RAGGED: Towards Informed Design of Scalable and Stable RAG Systems
von: Hsia, Jennifer, et al.
Veröffentlicht: (2024)
von: Hsia, Jennifer, et al.
Veröffentlicht: (2024)
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
von: Zhou, Xuhui, et al.
Veröffentlicht: (2026)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2026)
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
von: Wang, Xingyao, et al.
Veröffentlicht: (2025)
von: Wang, Xingyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025) -
TOM-SWE: User Mental Modeling For Software Engineering Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025) -
How Well Does Agent Development Reflect Real-World Work?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2026) -
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2026) -
Efficient On-Device Agents via Adaptive Context Management
von: Vijayvargiya, Sanidhya, et al.
Veröffentlicht: (2025)