AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Dongrui, Ren, Qihan, Qian, Chen, Shao, Shuai, Xie, Yuejin, Li, Yu, Yang, Zhonghao, Luo, Haoyu, Wang, Peng, Liu, Qingyu, Hu, Binxin, Tang, Ling, Mei, Jilin, Guo, Dadi, Yuan, Leitao, Yang, Junyao, Chen, Guanxu, Lin, Qihao, Yu, Yi, Zhang, Bo, Guo, Jiaxuan, Zhang, Jie, Shao, Wenqi, Deng, Huiqi, Xi, Zhiheng, Wang, Wenjie, Wang, Wenxuan, Shen, Wen, Chen, Zhikai, Xie, Haoyu, Tao, Jialing, Dai, Juntao, Ji, Jiaming, Ba, Zhongjie, Zhang, Linfeng, Liu, Yong, Zhang, Quanshi, Zhu, Lei, Wei, Zhihua, Xue, Hui, Lu, Chaochao, Shao, Jing, Hu, Xia |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
par: Liu, Dongrui, et autres
Publié: (2026)
par: Liu, Dongrui, et autres
Publié: (2026)
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
par: Chen, Guanxu, et autres
Publié: (2026)
par: Chen, Guanxu, et autres
Publié: (2026)
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
par: Shao, Shuai, et autres
Publié: (2025)
par: Shao, Shuai, et autres
Publié: (2025)
The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
par: Qian, Chen, et autres
Publié: (2026)
par: Qian, Chen, et autres
Publié: (2026)
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
par: Guo, Dadi, et autres
Publié: (2026)
par: Guo, Dadi, et autres
Publié: (2026)
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
par: Ren, Qihan, et autres
Publié: (2026)
par: Ren, Qihan, et autres
Publié: (2026)
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
par: Chen, Guanxu, et autres
Publié: (2025)
par: Chen, Guanxu, et autres
Publié: (2025)
Attributing Emergence in Million-Agent Systems
par: Tang, Ling, et autres
Publié: (2026)
par: Tang, Ling, et autres
Publié: (2026)
Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
par: Yang, Junyao, et autres
Publié: (2026)
par: Yang, Junyao, et autres
Publié: (2026)
Towards the Dynamics of a DNN Learning Symbolic Interactions
par: Ren, Qihan, et autres
Publié: (2024)
par: Ren, Qihan, et autres
Publié: (2024)
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
par: Chen, Guanxu, et autres
Publié: (2025)
par: Chen, Guanxu, et autres
Publié: (2025)
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report v1.5
par: Liu, Dongrui, et autres
Publié: (2026)
par: Liu, Dongrui, et autres
Publié: (2026)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
par: Lu, Xiaoya, et autres
Publié: (2025)
par: Lu, Xiaoya, et autres
Publié: (2025)
Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Step
par: Yang, Jingyi, et autres
Publié: (2025)
par: Yang, Jingyi, et autres
Publié: (2025)
Rethinking Entropy Regularization in Large Reasoning Models
par: Jiang, Yuxian, et autres
Publié: (2025)
par: Jiang, Yuxian, et autres
Publié: (2025)
ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
par: Yang, Junyao, et autres
Publié: (2026)
par: Yang, Junyao, et autres
Publié: (2026)
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents
par: Wang, Shengqin, et autres
Publié: (2026)
par: Wang, Shengqin, et autres
Publié: (2026)
ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards
par: Yan, Wentao, et autres
Publié: (2026)
par: Yan, Wentao, et autres
Publié: (2026)
Revisiting Generalization Power of a DNN in Terms of Symbolic Interactions
par: Cheng, Lei, et autres
Publié: (2025)
par: Cheng, Lei, et autres
Publié: (2025)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
par: Chen, Lu, et autres
Publié: (2024)
par: Chen, Lu, et autres
Publié: (2024)
LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint
par: Ma, Qianli, et autres
Publié: (2025)
par: Ma, Qianli, et autres
Publié: (2025)
The Better Angels of Machine Personality: How Personality Relates to LLM Safety
par: Zhang, Jie, et autres
Publié: (2024)
par: Zhang, Jie, et autres
Publié: (2024)
Are Your Agents Upward Deceivers?
par: Guo, Dadi, et autres
Publié: (2025)
par: Guo, Dadi, et autres
Publié: (2025)
Explaining Generalization Power of a DNN Using Interactive Concepts
par: Zhou, Huilin, et autres
Publié: (2023)
par: Zhou, Huilin, et autres
Publié: (2023)
Towards Attributions of Input Variables in a Coalition
par: Zheng, Xinhao, et autres
Publié: (2023)
par: Zheng, Xinhao, et autres
Publié: (2023)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
par: Li, Yu, et autres
Publié: (2026)
par: Li, Yu, et autres
Publié: (2026)
ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models
par: Wang, Fen, et autres
Publié: (2026)
par: Wang, Fen, et autres
Publié: (2026)
Understanding Generalization through Decision Pattern Shift
par: Deng, Huiqi, et autres
Publié: (2026)
par: Deng, Huiqi, et autres
Publié: (2026)
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
par: Guo, Dadi, et autres
Publié: (2025)
par: Guo, Dadi, et autres
Publié: (2025)
TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems
par: Xu, Chen, et autres
Publié: (2026)
par: Xu, Chen, et autres
Publié: (2026)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
par: Zhang, Boxuan, et autres
Publié: (2025)
par: Zhang, Boxuan, et autres
Publié: (2025)
REEF: Representation Encoding Fingerprints for Large Language Models
par: Zhang, Jie, et autres
Publié: (2024)
par: Zhang, Jie, et autres
Publié: (2024)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
par: Hu, Xuhao, et autres
Publié: (2024)
par: Hu, Xuhao, et autres
Publié: (2024)
Development and Preliminary Validation of an Immersive Virtual Reality Serious Game for Subway Fire Emergency Plan Training
par: Na Chen, et autres
Publié: (2026)
par: Na Chen, et autres
Publié: (2026)
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
par: Zhao, Haoyu, et autres
Publié: (2025)
par: Zhao, Haoyu, et autres
Publié: (2025)
SafePred: A Predictive Guardrail for Computer-Using Agents via World Models
par: Chen, Yurun, et autres
Publié: (2026)
par: Chen, Yurun, et autres
Publié: (2026)
Identifying Semantic Induction Heads to Understand In-Context Learning
par: Ren, Jie, et autres
Publié: (2024)
par: Ren, Jie, et autres
Publié: (2024)
Level-Navi Agent: A Framework and benchmark for Chinese Web Search Agents
par: Hu, Chuanrui, et autres
Publié: (2024)
par: Hu, Chuanrui, et autres
Publié: (2024)
EmpathyAgent: Can Embodied Agents Conduct Empathetic Actions?
par: Chen, Xinyan, et autres
Publié: (2025)
par: Chen, Xinyan, et autres
Publié: (2025)
A Prediction-as-Perception Framework for 3D Object Detection
par: Zhang, Song, et autres
Publié: (2026)
par: Zhang, Song, et autres
Publié: (2026)
Documents similaires
-
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
par: Liu, Dongrui, et autres
Publié: (2026) -
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
par: Chen, Guanxu, et autres
Publié: (2026) -
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
par: Shao, Shuai, et autres
Publié: (2025) -
The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
par: Qian, Chen, et autres
Publié: (2026) -
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
par: Guo, Dadi, et autres
Publié: (2026)