AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Dongrui, Li, Yu, Yang, Zhonghao, Wang, Peng, Chen, Guanxu, Xie, Yuejin, Mao, Qinghua, Qu, Wanying, Zhu, Yanxu, Zhou, Tianyi, Yuan, Leitao, Zheng, Zhijie, Lin, Qihao, Wang, Yimin, Luo, Haoyu, Shao, Shuai, Qian, Chen, Liu, Qingyu, Tang, Ling, Qin, Ruiyang, Ren, Qihan, Yang, Junxiao, Wang, Kun, Xi, Zhiheng, Zhang, Linfeng, Duan, Ranjie, Zhang, Bo, Wang, Wenjie, Shen, Wen, Zhang, Qiaosheng, Teng, Yan, Lu, Chaochao, Mei, Rui, Li, Man, Tao, Jialing, Lin, Xi, Zheng, Tianhang, Liu, Yong, Zhang, Quanshi, Zhu, Lei, Ma, Xingjun, Liu, Junhua, Xue, Hui, Zuo, Xiaoxiang, He, Xiangnan, Shen, Chao, Liu, Xianglong, Huang, Minlie, Shao, Jing, Hu, Xia |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
by: Liu, Dongrui, et al.
Published: (2026)
by: Liu, Dongrui, et al.
Published: (2026)
Attributing Emergence in Million-Agent Systems
by: Tang, Ling, et al.
Published: (2026)
by: Tang, Ling, et al.
Published: (2026)
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
by: Chen, Guanxu, et al.
Published: (2026)
by: Chen, Guanxu, et al.
Published: (2026)
PRISM: Preference-Aware Influence Function Based Data Selection Method for Efficient Fine-Tuning
by: Lin, Qihao, et al.
Published: (2026)
by: Lin, Qihao, et al.
Published: (2026)
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
by: Ren, Qihan, et al.
Published: (2026)
by: Ren, Qihan, et al.
Published: (2026)
Towards the Dynamics of a DNN Learning Symbolic Interactions
by: Ren, Qihan, et al.
Published: (2024)
by: Ren, Qihan, et al.
Published: (2024)
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025)
by: Shao, Shuai, et al.
Published: (2025)
Distributed Cooperative Optimization Control for Nonlinear Multi‐Agent Systems With Event‐Triggered Communication
by: Dan Liu, et al.
Published: (2026)
by: Dan Liu, et al.
Published: (2026)
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
by: Chen, Guanxu, et al.
Published: (2025)
by: Chen, Guanxu, et al.
Published: (2025)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
by: Yang, Jingyi, et al.
Published: (2025)
by: Yang, Jingyi, et al.
Published: (2025)
Where We Have Arrived in Proving the Emergence of Sparse Symbolic Concepts in AI Models
by: Ren, Qihan, et al.
Published: (2023)
by: Ren, Qihan, et al.
Published: (2023)
Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
by: Yang, Junyao, et al.
Published: (2026)
by: Yang, Junyao, et al.
Published: (2026)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
by: Yang, Zhonghao, et al.
Published: (2026)
by: Yang, Zhonghao, et al.
Published: (2026)
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
by: Zhang, Zhexin, et al.
Published: (2024)
by: Zhang, Zhexin, et al.
Published: (2024)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
by: Chen, Lu, et al.
Published: (2024)
by: Chen, Lu, et al.
Published: (2024)
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
by: Wang, Wenhan, et al.
Published: (2026)
by: Wang, Wenhan, et al.
Published: (2026)
Are Your Agents Upward Deceivers?
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
Independent characterization of the elastic and the mixing parts of hydrogel osmotic pressure
by: Shao, Zefan, et al.
Published: (2023)
by: Shao, Zefan, et al.
Published: (2023)
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
by: Chen, Guanxu, et al.
Published: (2025)
by: Chen, Guanxu, et al.
Published: (2025)
Rethinking Entropy Regularization in Large Reasoning Models
by: Jiang, Yuxian, et al.
Published: (2025)
by: Jiang, Yuxian, et al.
Published: (2025)
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
by: Guo, Dadi, et al.
Published: (2026)
by: Guo, Dadi, et al.
Published: (2026)
Explaining Generalization Power of a DNN Using Interactive Concepts
by: Zhou, Huilin, et al.
Published: (2023)
by: Zhou, Huilin, et al.
Published: (2023)
Model Structures Arising from Extendable Cotorsion Pairs
by: Shao, Qingyu, et al.
Published: (2025)
by: Shao, Qingyu, et al.
Published: (2025)
Superplatforms Have to Attack AI Agents
by: Lin, Jianghao, et al.
Published: (2025)
by: Lin, Jianghao, et al.
Published: (2025)
Revisiting Generalization Power of a DNN in Terms of Symbolic Interactions
by: Cheng, Lei, et al.
Published: (2025)
by: Cheng, Lei, et al.
Published: (2025)
Fast and Lightweight Novel View Synthesis with Differentiable Multiplane Image
by: Zhang, Kaidi, et al.
Published: (2026)
by: Zhang, Kaidi, et al.
Published: (2026)
AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent Systems
by: Chen, Yulang, et al.
Published: (2026)
by: Chen, Yulang, et al.
Published: (2026)
Finite Blocklength Covert Communication over Quasi-Static Multiple-Antenna Fading Channels
by: Liu, Changhong, et al.
Published: (2026)
by: Liu, Changhong, et al.
Published: (2026)
Identifying Semantic Induction Heads to Understand In-Context Learning
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
Fatigue Life Prediction and Reliability Assessment of CFRP Adhesively Bonded Joints in Offshore Wind Turbine Blade Applications: A Physics‐Informed Data‐Driven Approach
by: Zhenjiang Shao, et al.
Published: (2024)
by: Zhenjiang Shao, et al.
Published: (2024)
Screened d‐p Orbital Hybridization in Turing Structure of Confined Nickel for Sulfion Oxidation Accelerated Hydrogen Production
by: Yin Zhu, et al.
Published: (2024)
by: Yin Zhu, et al.
Published: (2024)
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report v1.5
by: Liu, Dongrui, et al.
Published: (2026)
by: Liu, Dongrui, et al.
Published: (2026)
TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
by: Yan, Lewen, et al.
Published: (2025)
by: Yan, Lewen, et al.
Published: (2025)
The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration
by: Wang, Zhexuan, et al.
Published: (2025)
by: Wang, Zhexuan, et al.
Published: (2025)
Few for Many: Tchebycheff Set Scalarization for Many-Objective Optimization
by: Lin, Xi, et al.
Published: (2024)
by: Lin, Xi, et al.
Published: (2024)
Cache Mechanism for Agent RAG Systems
by: Lin, Shuhang, et al.
Published: (2025)
by: Lin, Shuhang, et al.
Published: (2025)
Similar Items
-
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
by: Liu, Dongrui, et al.
Published: (2026) -
Attributing Emergence in Million-Agent Systems
by: Tang, Ling, et al.
Published: (2026) -
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
by: Chen, Guanxu, et al.
Published: (2026) -
PRISM: Preference-Aware Influence Function Based Data Selection Method for Efficient Fine-Tuning
by: Lin, Qihao, et al.
Published: (2026) -
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
by: Ren, Qihan, et al.
Published: (2026)