BraveGuard: From Open-World Threats to Safer Computer-Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yunhao, Ding, Yifan, Du, Xiaohu, Wen, Ming, Deng, Xinhao, Guo, Yanming, Xie, Yuxiang, Zheng, Baihui, Tan, Yingshui, Li, Yige, Wu, Yutao, Wang, Yixu, Cao, Kerui, Huang, Wenke, Ma, Xingjun, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025)
by: Li, Juncheng, et al.
Published: (2025)
RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting
by: Jiang, Yilei, et al.
Published: (2024)
by: Jiang, Yilei, et al.
Published: (2024)
Brave Newold World.
by: Rettig, James
Published: (1995)
by: Rettig, James
Published: (1995)
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
by: Yu, Jiaqi, et al.
Published: (2026)
by: Yu, Jiaqi, et al.
Published: (2026)
Information's Brave New World.
by: Malinconico, S. Michael
Published: (1992)
by: Malinconico, S. Michael
Published: (1992)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
by: Wu, Yutao, et al.
Published: (2025)
by: Wu, Yutao, et al.
Published: (2025)
A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
by: Ma, Xingjun, et al.
Published: (2026)
by: Ma, Xingjun, et al.
Published: (2026)
Smartcards in Libraries: A Brave New World.
by: Myhill, Martin
Published: (1998)
by: Myhill, Martin
Published: (1998)
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models
by: Tan, Yingshui, et al.
Published: (2024)
by: Tan, Yingshui, et al.
Published: (2024)
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
by: Zheng, Baihui, et al.
Published: (2025)
by: Zheng, Baihui, et al.
Published: (2025)
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
by: Zheng, Xiang, et al.
Published: (2026)
by: Zheng, Xiang, et al.
Published: (2026)
StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
Brave.Net.World: The Internet as a Disinformation Superhighway?
by: Floridi, Luciano
Published: (1996)
by: Floridi, Luciano
Published: (1996)
Brave Humanism
by: Godfrey, Mollie
Published: (2025)
by: Godfrey, Mollie
Published: (2025)
Ode to the Brave
Extracting Training Data from Unconditional Diffusion Models
by: Chen, Yunhao, et al.
Published: (2024)
by: Chen, Yunhao, et al.
Published: (2024)
Expose Before You Defend: Unifying and Enhancing Backdoor Defenses via Exposed Models
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
by: Cui, Kaiyuan, et al.
Published: (2026)
by: Cui, Kaiyuan, et al.
Published: (2026)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
by: Chen, Yunhao, et al.
Published: (2025)
by: Chen, Yunhao, et al.
Published: (2025)
A Brave New World: The Impact of Technology on Innovation Management
by: Dhruv Grewal, et al.
Published: (2025)
by: Dhruv Grewal, et al.
Published: (2025)
The Governance of Well-being: Towards a “Brave New World”?
by: Alex Romaní Rivera
Published: (2024)
by: Alex Romaní Rivera
Published: (2024)
fastml: Guarded Resampling Workflows for Safer Automated Machine Learning in R
by: Korkmaz, Selcuk, et al.
Published: (2026)
by: Korkmaz, Selcuk, et al.
Published: (2026)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026)
by: Zhao, Yunhan, et al.
Published: (2026)
Position: AI Safety Requires Effective Controllability
by: Li, Yige, et al.
Published: (2026)
by: Li, Yige, et al.
Published: (2026)
Braving Troubled Waters
by: van Ginkel, Rob
Published: (2010)
by: van Ginkel, Rob
Published: (2010)
Brave new family
Published: (1999)
Published: (1999)
Be Brave by Being Here
by: Julie Stivers
Published: (2023)
by: Julie Stivers
Published: (2023)
PurpCode: Reasoning for Safer Code Generation
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
New Liouville type theorems for the stationary Navier-Stokes equations
by: Tan, Wenke
Published: (2025)
by: Tan, Wenke
Published: (2025)
New Liouville type theorems for the stationary MHD equations in $\mathbb{R}^3$
by: Tan, Wenke
Published: (2025)
by: Tan, Wenke
Published: (2025)
Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats
by: Deng, Xinhao, et al.
Published: (2026)
by: Deng, Xinhao, et al.
Published: (2026)
New Technologies and Old-Fashioned Economics: Creating a Brave New World for U.S. Government Information Distribution and Use.
by: Smith, Diane
Published: (1999)
by: Smith, Diane
Published: (1999)
Engineering a Safer World
by: Leveson, Nancy G.
Published: (2019)
by: Leveson, Nancy G.
Published: (2019)
HoneypotNet: Backdoor Attacks Against Model Extraction
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
Brave Men / Ernie Pyle
by: Pyle, Ernie
Published: (1939)
by: Pyle, Ernie
Published: (1939)
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
by: Zhao, Yunhan, et al.
Published: (2024)
by: Zhao, Yunhan, et al.
Published: (2024)
Similar Items
-
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026) -
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
by: Feng, Yunhao, et al.
Published: (2026) -
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
by: Feng, Yunhao, et al.
Published: (2026) -
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025) -
RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting
by: Jiang, Yilei, et al.
Published: (2024)