ClawSafety: "Safe" LLMs, Unsafe Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wei, Bowen, Zhang, Yunbei, Pan, Jinhao, Mei, Kai, Wang, Xiao, Hamm, Jihun, Zhu, Ziwei, Ge, Yingqiang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Agents in the Wild: Safety, Society, and the Illusion of Sociality on Moltbook
par: Zhang, Yunbei, et autres
Publié: (2026)
par: Zhang, Yunbei, et autres
Publié: (2026)
Stop Comparing LLM Agents Without Disclosing the Harness
par: Zhang, Yunbei, et autres
Publié: (2026)
par: Zhang, Yunbei, et autres
Publié: (2026)
Visual Instance-aware Prompt Tuning
par: Xiao, Xi, et autres
Publié: (2025)
par: Xiao, Xi, et autres
Publié: (2025)
Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback
par: Wei, Bowen, et autres
Publié: (2026)
par: Wei, Bowen, et autres
Publié: (2026)
KnowBias: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
par: Pan, Jinhao, et autres
Publié: (2026)
par: Pan, Jinhao, et autres
Publié: (2026)
Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language Models
par: Cai, Wei, et autres
Publié: (2025)
par: Cai, Wei, et autres
Publié: (2025)
SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
par: Hao, Haochang, et autres
Publié: (2026)
par: Hao, Haochang, et autres
Publié: (2026)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
par: Liu, Songyang, et autres
Publié: (2026)
par: Liu, Songyang, et autres
Publié: (2026)
Promoting Online Safety by Simulating Unsafe Conversations with LLMs
par: Hoffman, Owen, et autres
Publié: (2025)
par: Hoffman, Owen, et autres
Publié: (2025)
Visual Exclusivity Attacks: Automatic Multimodal Red Teaming via Agentic Planning
par: Zhang, Yunbei, et autres
Publié: (2026)
par: Zhang, Yunbei, et autres
Publié: (2026)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
par: Jiang, Yukun, et autres
Publié: (2026)
par: Jiang, Yukun, et autres
Publié: (2026)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
par: Wang, Siyin, et autres
Publié: (2024)
par: Wang, Siyin, et autres
Publié: (2024)
OT-VP: Optimal Transport-guided Visual Prompting for Test-Time Adaptation
par: Zhang, Yunbei, et autres
Publié: (2024)
par: Zhang, Yunbei, et autres
Publié: (2024)
Understanding the Transferability of Representations via Task-Relatedness
par: Mehra, Akshay, et autres
Publié: (2023)
par: Mehra, Akshay, et autres
Publié: (2023)
Test-time Assessment of a Model's Performance on Unseen Domains via Optimal Transport
par: Mehra, Akshay, et autres
Publié: (2024)
par: Mehra, Akshay, et autres
Publié: (2024)
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
par: Li, Xiangyi, et autres
Publié: (2026)
par: Li, Xiangyi, et autres
Publié: (2026)
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
par: Hua, Wenyue, et autres
Publié: (2023)
par: Hua, Wenyue, et autres
Publié: (2023)
HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios
par: Pu, Jiayue, et autres
Publié: (2026)
par: Pu, Jiayue, et autres
Publié: (2026)
Interpretable Classification via a Rule Network with Selective Logical Operators
par: Wei, Bowen, et autres
Publié: (2024)
par: Wei, Bowen, et autres
Publié: (2024)
Advancing Interpretability in Text Classification through Prototype Learning
par: Wei, Bowen, et autres
Publié: (2024)
par: Wei, Bowen, et autres
Publié: (2024)
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
par: Ye, Bowen, et autres
Publié: (2026)
par: Ye, Bowen, et autres
Publié: (2026)
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
par: Zhang, Qiaohong, et autres
Publié: (2026)
par: Zhang, Qiaohong, et autres
Publié: (2026)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
par: Yuan, Youliang, et autres
Publié: (2024)
par: Yuan, Youliang, et autres
Publié: (2024)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
par: Yang, Zhonghao, et autres
Publié: (2026)
par: Yang, Zhonghao, et autres
Publié: (2026)
eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases
par: Wang, Janet, et autres
Publié: (2025)
par: Wang, Janet, et autres
Publié: (2025)
AIOS: LLM Agent Operating System
par: Mei, Kai, et autres
Publié: (2024)
par: Mei, Kai, et autres
Publié: (2024)
SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents
par: Cheng, Hao, et autres
Publié: (2026)
par: Cheng, Hao, et autres
Publié: (2026)
SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering
par: Zhu, Ningyan, et autres
Publié: (2026)
par: Zhu, Ningyan, et autres
Publié: (2026)
Safe-Support Q-Learning: Learning without Unsafe Exploration
par: Lim, Yeeun, et autres
Publié: (2026)
par: Lim, Yeeun, et autres
Publié: (2026)
MediaClaw: Multimodal Intelligent-Agent Platform Technical Report
par: Zhao, Shaoan, et autres
Publié: (2026)
par: Zhao, Shaoan, et autres
Publié: (2026)
Claw AI Lab: An Autonomous Multi-Agent Research Team
par: Wu, Fan, et autres
Publié: (2026)
par: Wu, Fan, et autres
Publié: (2026)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
par: Wang, Zijun, et autres
Publié: (2026)
par: Wang, Zijun, et autres
Publié: (2026)
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
par: Li, Xirui, et autres
Publié: (2026)
par: Li, Xirui, et autres
Publié: (2026)
ClawNet: Human-Symbiotic Agent Network for Cross-User Autonomous Cooperation
par: Yang, Zhiqin, et autres
Publié: (2026)
par: Yang, Zhiqin, et autres
Publié: (2026)
ClawLess: A Security Model of AI Agents
par: Lu, Hongyi, et autres
Publié: (2026)
par: Lu, Hongyi, et autres
Publié: (2026)
Subliminal Transfer of Unsafe Behaviors in AI Agent Distillation
par: Dang, Jacob, et autres
Publié: (2026)
par: Dang, Jacob, et autres
Publié: (2026)
ClawGym: A Scalable Framework for Building Effective Claw Agents
par: Bai, Fei, et autres
Publié: (2026)
par: Bai, Fei, et autres
Publié: (2026)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
par: Zhang, Yuxuan, et autres
Publié: (2026)
par: Zhang, Yuxuan, et autres
Publié: (2026)
TrafficClaw: A Generalizable LLM Agent in the Unified Physical Environment for Urban Traffic Control
par: Lai, Siqi, et autres
Publié: (2026)
par: Lai, Siqi, et autres
Publié: (2026)
MolClaw: An Autonomous Agent with Hierarchical Skills for Drug Molecule Evaluation, Screening, and Optimization
par: Zhang, Lisheng, et autres
Publié: (2026)
par: Zhang, Lisheng, et autres
Publié: (2026)
Documents similaires
-
Agents in the Wild: Safety, Society, and the Illusion of Sociality on Moltbook
par: Zhang, Yunbei, et autres
Publié: (2026) -
Stop Comparing LLM Agents Without Disclosing the Harness
par: Zhang, Yunbei, et autres
Publié: (2026) -
Visual Instance-aware Prompt Tuning
par: Xiao, Xi, et autres
Publié: (2025) -
Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback
par: Wei, Bowen, et autres
Publié: (2026) -
KnowBias: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
par: Pan, Jinhao, et autres
Publié: (2026)