OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Kuntz, Thomas, Duzan, Agatha, Zhao, Hao, Croce, Francesco, Kolter, Zico, Flammarion, Nicolas, Andriushchenko, Maksym |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026)
by: Rank, Ben, et al.
Published: (2026)
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
by: Xiao, Yijia, et al.
Published: (2025)
by: Xiao, Yijia, et al.
Published: (2025)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
by: Raghavendra, Mohit, et al.
Published: (2026)
by: Raghavendra, Mohit, et al.
Published: (2026)
Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents
by: Cai, Yuandao, et al.
Published: (2026)
by: Cai, Yuandao, et al.
Published: (2026)
Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning
by: Golubev, Alexander, et al.
Published: (2025)
by: Golubev, Alexander, et al.
Published: (2025)
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
by: Wei, Jinbiao, et al.
Published: (2026)
by: Wei, Jinbiao, et al.
Published: (2026)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
by: LeVine, Will, et al.
Published: (2026)
by: LeVine, Will, et al.
Published: (2026)
Why Do We Need Weight Decay in Modern Deep Learning?
by: D'Angelo, Francesco, et al.
Published: (2023)
by: D'Angelo, Francesco, et al.
Published: (2023)
Automating Benchmark Design
by: Dsouza, Amanda, et al.
Published: (2025)
by: Dsouza, Amanda, et al.
Published: (2025)
AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development
by: Agarwal, Shyam, et al.
Published: (2026)
by: Agarwal, Shyam, et al.
Published: (2026)
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
by: Feng, Yunfei, et al.
Published: (2026)
by: Feng, Yunfei, et al.
Published: (2026)
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
by: Yu, Shasha, et al.
Published: (2026)
by: Yu, Shasha, et al.
Published: (2026)
SMARLA: A Safety Monitoring Approach for Deep Reinforcement Learning Agents
by: Zolfagharian, Amirhossein, et al.
Published: (2023)
by: Zolfagharian, Amirhossein, et al.
Published: (2023)
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
by: Hu, Qi, et al.
Published: (2026)
by: Hu, Qi, et al.
Published: (2026)
Unified Implementations of Recurrent Neural Networks in Multiple Deep Learning Frameworks
by: Martinuzzi, Francesco
Published: (2025)
by: Martinuzzi, Francesco
Published: (2025)
Uncertainty-Guided Label Rebalancing for CPS Safety Monitoring
by: Ayotunde, John, et al.
Published: (2026)
by: Ayotunde, John, et al.
Published: (2026)
Reconciling Safety Measurement and Dynamic Assurance
by: Denney, Ewen, et al.
Published: (2024)
by: Denney, Ewen, et al.
Published: (2024)
Architecting software monitors for control-flow anomaly detection through large language models and conformance checking
by: Vitale, Francesco, et al.
Published: (2025)
by: Vitale, Francesco, et al.
Published: (2025)
Evaluating Reinforcement Learning Safety and Trustworthiness in Cyber-Physical Systems
by: Dearstyne, Katherine, et al.
Published: (2025)
by: Dearstyne, Katherine, et al.
Published: (2025)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
Concept-Guided LLM Agents for Human-AI Safety Codesign
by: Geissler, Florian, et al.
Published: (2024)
by: Geissler, Florian, et al.
Published: (2024)
Measuring Agents in Production
by: Pan, Melissa Z., et al.
Published: (2025)
by: Pan, Melissa Z., et al.
Published: (2025)
Towards an Argument Pattern for the Use of Safety Performance Indicators
by: Ratiu, Daniel, et al.
Published: (2024)
by: Ratiu, Daniel, et al.
Published: (2024)
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
by: Jha, Smriti, et al.
Published: (2026)
by: Jha, Smriti, et al.
Published: (2026)
OSS-Bench: Benchmark Generator for Coding LLMs
by: Jiang, Yuancheng, et al.
Published: (2025)
by: Jiang, Yuancheng, et al.
Published: (2025)
Supporting Safety Analysis of Image-processing DNNs through Clustering-based Approaches
by: Attaoui, Mohammed Oualid, et al.
Published: (2023)
by: Attaoui, Mohammed Oualid, et al.
Published: (2023)
AgentForge: A Flexible Low-Code Platform for Reinforcement Learning Agent Design
by: Junior, Francisco Erivaldo Fernandes, et al.
Published: (2024)
by: Junior, Francisco Erivaldo Fernandes, et al.
Published: (2024)
Assessing the Use of AutoML for Data-Driven Software Engineering
by: Calefato, Fabio, et al.
Published: (2023)
by: Calefato, Fabio, et al.
Published: (2023)
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025)
by: Eliseeva, Aleksandra, et al.
Published: (2025)
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
by: Glukhov, Evgeniy, et al.
Published: (2025)
by: Glukhov, Evgeniy, et al.
Published: (2025)
ThrowBench: Benchmarking LLMs by Predicting Runtime Exceptions
by: Prenner, Julian Aron, et al.
Published: (2025)
by: Prenner, Julian Aron, et al.
Published: (2025)
Exploring Code Language Models for Automated HLS-based Hardware Generation: Benchmark, Infrastructure and Analysis
by: Gai, Jiahao, et al.
Published: (2025)
by: Gai, Jiahao, et al.
Published: (2025)
FROAV: A Framework for RAG Observation and Agent Verification -- Lowering the Barrier to LLM Agent Research
by: Lin, Tzu-Hsuan, et al.
Published: (2026)
by: Lin, Tzu-Hsuan, et al.
Published: (2026)
UCRBench: Benchmarking LLMs on Use Case Recovery
by: Xiao, Shuyuan, et al.
Published: (2025)
by: Xiao, Shuyuan, et al.
Published: (2025)
Similar Items
-
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024) -
Is In-Context Learning Sufficient for Instruction Following in LLMs?
by: Zhao, Hao, et al.
Published: (2024) -
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
by: Rank, Ben, et al.
Published: (2026) -
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning
by: Zhao, Hao, et al.
Published: (2024) -
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)