SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Xueyang, Wang, Weidong, Lu, Lin, Shi, Jiawen, Tie, Guiyao, Xu, Yongtian, Chen, Lixing, Zhou, Pan, Gong, Neil Zhenqiang, Sun, Lichao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization
by: Zhou, Xueyang, et al.
Published: (2025)
by: Zhou, Xueyang, et al.
Published: (2025)
Step-Aware Residual-Guided Diffusion for EEG Spatial Super-Resolution
by: Liu, Hongjun, et al.
Published: (2025)
by: Liu, Hongjun, et al.
Published: (2025)
Territory Paint Wars: Diagnosing and Mitigating Failure Modes in Competitive Multi-Agent PPO
by: Singh, Diyansha
Published: (2026)
by: Singh, Diyansha
Published: (2026)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023)
by: Brännvall, Rickard, et al.
Published: (2023)
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
by: Yamchote, Phaphontee, et al.
Published: (2025)
by: Yamchote, Phaphontee, et al.
Published: (2025)
Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols
by: Reitich, Fernando
Published: (2026)
by: Reitich, Fernando
Published: (2026)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025)
by: Tiwari, Dhruv
Published: (2025)
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026)
by: Zhou, Qinghua, et al.
Published: (2026)
Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Evaluating Model-Agnostic Meta-Learning on MetaWorld ML10 Benchmark: Fast Adaptation in Robotic Manipulation Tasks
by: Atamuradov, Sanjar
Published: (2025)
by: Atamuradov, Sanjar
Published: (2025)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Distinguished In Uniform: Self Attention Vs. Virtual Nodes
by: Rosenbluth, Eran, et al.
Published: (2024)
by: Rosenbluth, Eran, et al.
Published: (2024)
Deep Hedging Under Non-Convexity: Limitations and a Case for AlphaZero
by: Maggiolo, Matteo, et al.
Published: (2025)
by: Maggiolo, Matteo, et al.
Published: (2025)
Model Capability Dominates: Inference-Time Optimization Lessons from AIMO 3
by: Nitarach, Natapong
Published: (2026)
by: Nitarach, Natapong
Published: (2026)
Stage-wise Dynamics of Classifier-Free Guidance in Diffusion Models
by: Jin, Cheng, et al.
Published: (2025)
by: Jin, Cheng, et al.
Published: (2025)
Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems
by: Parris, William
Published: (2026)
by: Parris, William
Published: (2026)
Drift-Resilient TabPFN: In-Context Learning Temporal Distribution Shifts on Tabular Data
by: Helli, Kai, et al.
Published: (2024)
by: Helli, Kai, et al.
Published: (2024)
FeNeC: Enhancing Continual Learning via Feature Clustering with Neighbor- or Logit-Based Classification
by: Książek, Kamil, et al.
Published: (2025)
by: Książek, Kamil, et al.
Published: (2025)
X-Factor: Quality Is a Dataset-Intrinsic Property
by: Couch, Josiah, et al.
Published: (2025)
by: Couch, Josiah, et al.
Published: (2025)
A Constraint-Preserving Neural Network Approach for Solving Mean-Field Games Equilibrium
by: Liu, Jinwei, et al.
Published: (2025)
by: Liu, Jinwei, et al.
Published: (2025)
The Domain Mixed Unit: A New Neural Arithmetic Layer
by: Curry, Paul
Published: (2025)
by: Curry, Paul
Published: (2025)
Deep learning four decades of human migration
by: Gaskin, Thomas, et al.
Published: (2025)
by: Gaskin, Thomas, et al.
Published: (2025)
IntSeqBERT: Learning Arithmetic Structure in OEIS via Modulo-Spectrum Embeddings
by: Nakasho, Kazuhisa
Published: (2026)
by: Nakasho, Kazuhisa
Published: (2026)
Your contrastive learning problem is secretly a distribution alignment problem
by: Chen, Zihao, et al.
Published: (2025)
by: Chen, Zihao, et al.
Published: (2025)
Developing Explainable Machine Learning Model using Augmented Concept Activation Vector
by: Hassanpour, Reza, et al.
Published: (2024)
by: Hassanpour, Reza, et al.
Published: (2024)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
by: Wang, Zhen, et al.
Published: (2025)
by: Wang, Zhen, et al.
Published: (2025)
Just In Time Transformers
by: Benali, Ahmed Ala Eddine, et al.
Published: (2024)
by: Benali, Ahmed Ala Eddine, et al.
Published: (2024)
Analyzing Closed-loop Training Techniques for Realistic Traffic Agent Models in Autonomous Highway Driving Simulations
by: Bitzer, Matthias, et al.
Published: (2024)
by: Bitzer, Matthias, et al.
Published: (2024)
Improving LLM Agent Planning with In-Context Learning via Atomic Fact Augmentation and Lookahead Search
by: Holt, Samuel, et al.
Published: (2025)
by: Holt, Samuel, et al.
Published: (2025)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
by: Pepe, Alberto, et al.
Published: (2026)
by: Pepe, Alberto, et al.
Published: (2026)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026)
by: Saini, Saurabh, et al.
Published: (2026)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
by: Lasbordes, Maxence, et al.
Published: (2026)
by: Lasbordes, Maxence, et al.
Published: (2026)
ProactBench: Beyond What The User Asked For
by: Harfi, Sepehr, et al.
Published: (2026)
by: Harfi, Sepehr, et al.
Published: (2026)
On the Origin of Algorithmic Progress in AI
by: Gundlach, Hans, et al.
Published: (2025)
by: Gundlach, Hans, et al.
Published: (2025)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
by: Lee, Wooin, et al.
Published: (2026)
by: Lee, Wooin, et al.
Published: (2026)
Advances in Set Function Learning: A Survey of Techniques and Applications
by: Xie, Jiahao, et al.
Published: (2025)
by: Xie, Jiahao, et al.
Published: (2025)
Similar Items
-
BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization
by: Zhou, Xueyang, et al.
Published: (2025) -
Step-Aware Residual-Guided Diffusion for EEG Spatial Super-Resolution
by: Liu, Hongjun, et al.
Published: (2025) -
Territory Paint Wars: Diagnosing and Mitigating Failure Modes in Competitive Multi-Agent PPO
by: Singh, Diyansha
Published: (2026) -
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023) -
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
by: Yamchote, Phaphontee, et al.
Published: (2025)