MirrorGuard: Toward Secure Computer-Use Agents via Simulation-to-Real Reasoning Correction
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Wenqi, Shen, Yulin, Jiang, Changyue, Dai, Jiarun, Hong, Geng, Pan, Xudong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents
di: Wu, Xinyi, et al.
Pubblicazione: (2026)
di: Wu, Xinyi, et al.
Pubblicazione: (2026)
AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent
di: Luo, Jiaqi, et al.
Pubblicazione: (2026)
di: Luo, Jiaqi, et al.
Pubblicazione: (2026)
Invisible Threats from Model Context Protocol: Generating Stealthy Injection Payload via Tree-based Adaptive Search
di: Shen, Yulin, et al.
Pubblicazione: (2026)
di: Shen, Yulin, et al.
Pubblicazione: (2026)
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
di: Fan, Yihe, et al.
Pubblicazione: (2026)
di: Fan, Yihe, et al.
Pubblicazione: (2026)
When Bots Take the Bait: Exposing and Mitigating the Emerging Social Engineering Attack in Web Automation Agent
di: Wu, Xinyi, et al.
Pubblicazione: (2026)
di: Wu, Xinyi, et al.
Pubblicazione: (2026)
OpenDeception: Learning Deception and Trust in Human-AI Interaction via Multi-Agent Simulation
di: Wu, Yichen, et al.
Pubblicazione: (2025)
di: Wu, Yichen, et al.
Pubblicazione: (2025)
Frontier AI systems have surpassed the self-replicating red line
di: Pan, Xudong, et al.
Pubblicazione: (2024)
di: Pan, Xudong, et al.
Pubblicazione: (2024)
GuardReasoner: Towards Reasoning-based LLM Safeguards
di: Liu, Yue, et al.
Pubblicazione: (2025)
di: Liu, Yue, et al.
Pubblicazione: (2025)
FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding
di: Yang, Jinghan, et al.
Pubblicazione: (2026)
di: Yang, Jinghan, et al.
Pubblicazione: (2026)
ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models
di: Li, Changyi, et al.
Pubblicazione: (2025)
di: Li, Changyi, et al.
Pubblicazione: (2025)
Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
di: Mai, Wuyuao, et al.
Pubblicazione: (2025)
di: Mai, Wuyuao, et al.
Pubblicazione: (2025)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
di: Liu, Yue, et al.
Pubblicazione: (2025)
di: Liu, Yue, et al.
Pubblicazione: (2025)
Secure and Efficient Access Control for Computer-Use Agents via Context Space
di: Gong, Haochen, et al.
Pubblicazione: (2025)
di: Gong, Haochen, et al.
Pubblicazione: (2025)
ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
di: Wang, Yuquan, et al.
Pubblicazione: (2025)
di: Wang, Yuquan, et al.
Pubblicazione: (2025)
BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
di: Miao, Rui, et al.
Pubblicazione: (2025)
di: Miao, Rui, et al.
Pubblicazione: (2025)
Large language model-powered AI systems achieve self-replication with no human intervention
di: Pan, Xudong, et al.
Pubblicazione: (2025)
di: Pan, Xudong, et al.
Pubblicazione: (2025)
Feedback-Guided Extraction of Knowledge Base from Retrieval-Augmented LLM Applications
di: Jiang, Changyue, et al.
Pubblicazione: (2024)
di: Jiang, Changyue, et al.
Pubblicazione: (2024)
CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents
di: Foerster, Hanna, et al.
Pubblicazione: (2026)
di: Foerster, Hanna, et al.
Pubblicazione: (2026)
Toward Agents That Reason About Their Computation
di: Orenstein, Adrian, et al.
Pubblicazione: (2025)
di: Orenstein, Adrian, et al.
Pubblicazione: (2025)
AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
di: Zhang, Xuan, et al.
Pubblicazione: (2025)
di: Zhang, Xuan, et al.
Pubblicazione: (2025)
Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems
di: Fan, Yihe, et al.
Pubblicazione: (2025)
di: Fan, Yihe, et al.
Pubblicazione: (2025)
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
di: Pan, Rui, et al.
Pubblicazione: (2025)
di: Pan, Rui, et al.
Pubblicazione: (2025)
MCPZoo: A Large-Scale Dataset of Runnable Model Context Protocol Servers for AI Agent
di: Wu, Mengying, et al.
Pubblicazione: (2025)
di: Wu, Mengying, et al.
Pubblicazione: (2025)
When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
di: Zhou, Xiaolin, et al.
Pubblicazione: (2026)
di: Zhou, Xiaolin, et al.
Pubblicazione: (2026)
X-Guard: Multilingual Guard Agent for Content Moderation
di: Upadhayay, Bibek, et al.
Pubblicazione: (2025)
di: Upadhayay, Bibek, et al.
Pubblicazione: (2025)
Cognitive Duality for Adaptive Web Agents
di: Liu, Jiarun, et al.
Pubblicazione: (2025)
di: Liu, Jiarun, et al.
Pubblicazione: (2025)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
di: Liu, Wenrui, et al.
Pubblicazione: (2025)
di: Liu, Wenrui, et al.
Pubblicazione: (2025)
Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detection
di: Feng, Yang, et al.
Pubblicazione: (2025)
di: Feng, Yang, et al.
Pubblicazione: (2025)
StruPhantom: Evolutionary Injection Attacks on Black-Box Tabular Agents Powered by Large Language Models
di: Feng, Yang, et al.
Pubblicazione: (2025)
di: Feng, Yang, et al.
Pubblicazione: (2025)
Decoupling Reasoning and Knowledge Injection for In-Context Knowledge Editing
di: Wang, Changyue, et al.
Pubblicazione: (2025)
di: Wang, Changyue, et al.
Pubblicazione: (2025)
AgentGuard: Runtime Verification of AI Agents
di: Koohestani, Roham
Pubblicazione: (2025)
di: Koohestani, Roham
Pubblicazione: (2025)
On the Reliability of Computer Use Agents
di: Gonzalez-Pumariega, Gonzalo, et al.
Pubblicazione: (2026)
di: Gonzalez-Pumariega, Gonzalo, et al.
Pubblicazione: (2026)
HarmonyGuard: Toward Safety and Utility in Web Agents via Adaptive Policy Enhancement and Dual-Objective Optimization
di: Chen, Yurun, et al.
Pubblicazione: (2025)
di: Chen, Yurun, et al.
Pubblicazione: (2025)
MirrorShield: Towards Universal Defense Against Jailbreaks via Entropy-Guided Mirror Crafting
di: Pu, Rui, et al.
Pubblicazione: (2025)
di: Pu, Rui, et al.
Pubblicazione: (2025)
Self-Guard: Defending Large Reasoning Models via enhanced self-reflection
di: Zheng, Jingnan, et al.
Pubblicazione: (2026)
di: Zheng, Jingnan, et al.
Pubblicazione: (2026)
$R^2$-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning
di: Kang, Mintong, et al.
Pubblicazione: (2024)
di: Kang, Mintong, et al.
Pubblicazione: (2024)
GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
di: Zhuang, Yuchen, et al.
Pubblicazione: (2025)
di: Zhuang, Yuchen, et al.
Pubblicazione: (2025)
Web-CogReasoner: Towards Knowledge-Induced Cognitive Reasoning for Web Agents
di: Guo, Yuhan, et al.
Pubblicazione: (2025)
di: Guo, Yuhan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents
di: Wu, Xinyi, et al.
Pubblicazione: (2026) -
AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent
di: Luo, Jiaqi, et al.
Pubblicazione: (2026) -
Invisible Threats from Model Context Protocol: Generating Stealthy Injection Payload via Tree-based Adaptive Search
di: Shen, Yulin, et al.
Pubblicazione: (2026) -
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
di: Fan, Yihe, et al.
Pubblicazione: (2026) -
When Bots Take the Bait: Exposing and Mitigating the Emerging Social Engineering Attack in Web Automation Agent
di: Wu, Xinyi, et al.
Pubblicazione: (2026)