The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Xuwei, Zhai, Skylar, Song, Linxin, Li, Jiate, Shi, Taiwei, Meade, Nicholas, Reddy, Siva, Kang, Jian, Zhao, Jieyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video-Based Reward Modeling for Computer-Use Agents
by: Song, Linxin, et al.
Published: (2026)
by: Song, Linxin, et al.
Published: (2026)
The Hallucination Tax of Reinforcement Finetuning
by: Song, Linxin, et al.
Published: (2025)
by: Song, Linxin, et al.
Published: (2025)
Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base
by: Song, Linxin, et al.
Published: (2025)
by: Song, Linxin, et al.
Published: (2025)
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
by: BehnamGhader, Parishad, et al.
Published: (2025)
by: BehnamGhader, Parishad, et al.
Published: (2025)
Investigating Adversarial Trigger Transfer in Large Language Models
by: Meade, Nicholas, et al.
Published: (2024)
by: Meade, Nicholas, et al.
Published: (2024)
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
by: Shi, Taiwei, et al.
Published: (2025)
by: Shi, Taiwei, et al.
Published: (2025)
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
by: Adlakha, Vaibhav, et al.
Published: (2023)
by: Adlakha, Vaibhav, et al.
Published: (2023)
Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
by: Jiang, Linxi, et al.
Published: (2026)
by: Jiang, Linxi, et al.
Published: (2026)
CoAct-1: Computer-using Multi-Agent System with Coding Actions
by: Song, Linxin, et al.
Published: (2025)
by: Song, Linxin, et al.
Published: (2025)
Experiential Reinforcement Learning
by: Shi, Taiwei, et al.
Published: (2026)
by: Shi, Taiwei, et al.
Published: (2026)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
Safer-Instruct: Aligning Language Models with Automated Preference Data
by: Shi, Taiwei, et al.
Published: (2023)
by: Shi, Taiwei, et al.
Published: (2023)
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
by: Yang, Tiankai, et al.
Published: (2026)
by: Yang, Tiankai, et al.
Published: (2026)
Structured Distillation of Web Agent Capabilities Enables Generalization
by: Lù, Xing Han, et al.
Published: (2026)
by: Lù, Xing Han, et al.
Published: (2026)
The Bandit's Blind Spot: The Critical Role of User State Representation in Recommender Systems
by: Pires, Pedro R., et al.
Published: (2026)
by: Pires, Pedro R., et al.
Published: (2026)
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
by: Xiong, Yuan, et al.
Published: (2025)
by: Xiong, Yuan, et al.
Published: (2025)
Auditable Agents
by: Nian, Yi, et al.
Published: (2026)
by: Nian, Yi, et al.
Published: (2026)
Offline Training of Language Model Agents with Functions as Learnable Weights
by: Zhang, Shaokun, et al.
Published: (2024)
by: Zhang, Shaokun, et al.
Published: (2024)
LLMs-based Few-Shot Disease Predictions using EHR: A Novel Approach Combining Predictive Agent Reasoning and Critical Agent Instruction
by: Cui, Hejie, et al.
Published: (2024)
by: Cui, Hejie, et al.
Published: (2024)
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
by: Pai, Aaditya
Published: (2026)
by: Pai, Aaditya
Published: (2026)
Adaptive In-conversation Team Building for Language Model Agents
by: Song, Linxin, et al.
Published: (2024)
by: Song, Linxin, et al.
Published: (2024)
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
by: Zhai, Skylar, et al.
Published: (2026)
by: Zhai, Skylar, et al.
Published: (2026)
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
by: Li, Xueyi, et al.
Published: (2026)
by: Li, Xueyi, et al.
Published: (2026)
Illuminating Blind Spots of Language Models with Targeted Agent-in-the-Loop Synthetic Data
by: Lippmann, Philip, et al.
Published: (2024)
by: Lippmann, Philip, et al.
Published: (2024)
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation
by: Pan, Yijun, et al.
Published: (2025)
by: Pan, Yijun, et al.
Published: (2025)
LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
by: Chen, Tianyu, et al.
Published: (2026)
by: Chen, Tianyu, et al.
Published: (2026)
OrchVis: Hierarchical Multi-Agent Orchestration for Human Oversight
by: Zhou, Jieyu
Published: (2025)
by: Zhou, Jieyu
Published: (2025)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
by: Nawal, Aditya, et al.
Published: (2026)
by: Nawal, Aditya, et al.
Published: (2026)
Blind Spots
by: Drucker, Johanna
Published: (2009)
by: Drucker, Johanna
Published: (2009)
Disentangling Electronic and Phononic Thermal Transport Across 2D Interfaces
by: Zhai, Linxin, et al.
Published: (2024)
by: Zhai, Linxin, et al.
Published: (2024)
Algorithmic Decision-Making under Agents with Persistent Improvement
by: Xie, Tian, et al.
Published: (2024)
by: Xie, Tian, et al.
Published: (2024)
Overcoming Blind Spots: Occlusion Considerations for Improved Autonomous Driving Safety
by: Moller, Korbinian, et al.
Published: (2024)
by: Moller, Korbinian, et al.
Published: (2024)
Lost in Instructions: Study of Blind Users' Experiences with DIY Manuals and AI-Rewritten Instructions for Assembly, Operation, and Troubleshooting of Tangible Products
by: Reddy, Monalika Padma, et al.
Published: (2026)
by: Reddy, Monalika Padma, et al.
Published: (2026)
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning
by: Zhang, Le, et al.
Published: (2025)
by: Zhang, Le, et al.
Published: (2025)
Mind the Blind Spot
by: Kelly Singleton, et al.
Published: (2026)
by: Kelly Singleton, et al.
Published: (2026)
A Systematization of Security Vulnerabilities in Computer Use Agents
by: Jones, Daniel, et al.
Published: (2025)
by: Jones, Daniel, et al.
Published: (2025)
From Defense to Advocacy: Empowering Users to Leverage the Blind Spot of AI Inference
by: Wei, Yumou, et al.
Published: (2026)
by: Wei, Yumou, et al.
Published: (2026)
How to Get Your LLM to Generate Challenging Problems for Evaluation
by: Patel, Arkil, et al.
Published: (2025)
by: Patel, Arkil, et al.
Published: (2025)
Similar Items
-
Video-Based Reward Modeling for Computer-Use Agents
by: Song, Linxin, et al.
Published: (2026) -
The Hallucination Tax of Reinforcement Finetuning
by: Song, Linxin, et al.
Published: (2025) -
Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base
by: Song, Linxin, et al.
Published: (2025) -
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
by: BehnamGhader, Parishad, et al.
Published: (2025) -
Investigating Adversarial Trigger Transfer in Large Language Models
by: Meade, Nicholas, et al.
Published: (2024)