When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Jones, Jaylen, Zhang, Zhehao, Ning, Yuting, Fosler-Lussier, Eric, St-Charles, Pierre-Luc, Bengio, Yoshua, Song, Dawn, Su, Yu, Sun, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
by: Jones, Jaylen, et al.
Published: (2024)
by: Jones, Jaylen, et al.
Published: (2024)
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
by: Liao, Zeyi, et al.
Published: (2025)
by: Liao, Zeyi, et al.
Published: (2025)
When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents
by: Ning, Yuting, et al.
Published: (2026)
by: Ning, Yuting, et al.
Published: (2026)
Music on the Move
by: Fosler-Lussier, Danielle
Published: (2020)
by: Fosler-Lussier, Danielle
Published: (2020)
Active Attacks: Red-teaming LLMs via Adaptive Environments
by: Yun, Taeyoung, et al.
Published: (2025)
by: Yun, Taeyoung, et al.
Published: (2025)
Improving Transducer-Based Spoken Language Understanding with Self-Conditioned CTC and Knowledge Transfer
by: Sunder, Vishal, et al.
Published: (2025)
by: Sunder, Vishal, et al.
Published: (2025)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers
by: Serai, Prashant, et al.
Published: (2024)
by: Serai, Prashant, et al.
Published: (2024)
End-to-End Diarization utilizing Attractor Deep Clustering
by: Palzer, David, et al.
Published: (2025)
by: Palzer, David, et al.
Published: (2025)
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
by: Palzer, David, et al.
Published: (2025)
by: Palzer, David, et al.
Published: (2025)
Machine learning and information theory concepts towards an AI Mathematician
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts
by: Kumar, Vishal, et al.
Published: (2024)
by: Kumar, Vishal, et al.
Published: (2024)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
by: Wu, Yixin, et al.
Published: (2023)
by: Wu, Yixin, et al.
Published: (2023)
Baking Symmetry into GFlowNets
by: Ma, George, et al.
Published: (2024)
by: Ma, George, et al.
Published: (2024)
VISTA: Verification In Sequential Turn-based Assessment
by: Lewis, Ashley, et al.
Published: (2025)
by: Lewis, Ashley, et al.
Published: (2025)
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances
by: Chun, Jiyun, et al.
Published: (2026)
by: Chun, Jiyun, et al.
Published: (2026)
Can a Bayesian Oracle Prevent Harm from an Agent?
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
by: Ginjala, Srishti, et al.
Published: (2026)
by: Ginjala, Srishti, et al.
Published: (2026)
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
by: Jiralerspong, Thomas, et al.
Published: (2026)
by: Jiralerspong, Thomas, et al.
Published: (2026)
Hijack-GAN: Unintended-Use of Pretrained, Black-Box GANs
by: Wang, Hui-Po, et al.
Published: (2020)
by: Wang, Hui-Po, et al.
Published: (2020)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)
by: Kaunismaa, Jackson, et al.
Published: (2026)
Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs
by: Yu, Mingyu, et al.
Published: (2026)
by: Yu, Mingyu, et al.
Published: (2026)
The Surprising Harmfulness of Benign Overfitting for Adversarial Robustness
by: Hao, Yifan, et al.
Published: (2024)
by: Hao, Yifan, et al.
Published: (2024)
When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
by: Dingeto, Hiskias, et al.
Published: (2025)
by: Dingeto, Hiskias, et al.
Published: (2025)
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
On Generalization for Generative Flow Networks
by: Krichel, Anas, et al.
Published: (2024)
by: Krichel, Anas, et al.
Published: (2024)
Interventional Causal Representation Learning
by: Ahuja, Kartik, et al.
Published: (2022)
by: Ahuja, Kartik, et al.
Published: (2022)
A Complexity-Based Theory of Compositionality
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
Visual symbolic mechanisms: Emergent symbol processing in vision language models
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
Relative Trajectory Balance is equivalent to Trust-PCL
by: Deleu, Tristan, et al.
Published: (2025)
by: Deleu, Tristan, et al.
Published: (2025)
Fast Monte Carlo Tree Diffusion: 100x Speedup via Parallel Sparse Planning
by: Yoon, Jaesik, et al.
Published: (2025)
by: Yoon, Jaesik, et al.
Published: (2025)
In-Context Parametric Inference: Point or Distribution Estimators?
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
Balancing Efficiency and Safety: How and When Algorithmic Management Induces Gig Workers' Unsafe Behavior
by: Yanghao Zhu, et al.
Published: (2026)
by: Yanghao Zhu, et al.
Published: (2026)
GFlowNet Foundations
by: Bengio, Yoshua, et al.
Published: (2021)
by: Bengio, Yoshua, et al.
Published: (2021)
When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks
by: Lo, Chung-Hsiang, et al.
Published: (2026)
by: Lo, Chung-Hsiang, et al.
Published: (2026)
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
by: Williams-King, David, et al.
Published: (2025)
by: Williams-King, David, et al.
Published: (2025)
RL, but don't do anything I wouldn't do
by: Cohen, Michael K., et al.
Published: (2024)
by: Cohen, Michael K., et al.
Published: (2024)
LLMGuard: Guarding Against Unsafe LLM Behavior
by: Goyal, Shubh, et al.
Published: (2024)
by: Goyal, Shubh, et al.
Published: (2024)
Similar Items
-
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
by: Jones, Jaylen, et al.
Published: (2024) -
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
by: Liao, Zeyi, et al.
Published: (2025) -
When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents
by: Ning, Yuting, et al.
Published: (2026) -
Music on the Move
by: Fosler-Lussier, Danielle
Published: (2020) -
Active Attacks: Red-teaming LLMs via Adaptive Environments
by: Yun, Taeyoung, et al.
Published: (2025)