ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yu, Luo, Haoyu, Xie, Yuejin, Fu, Yuqian, Yang, Zhonghao, Shao, Shuai, Ren, Qihan, Qu, Wanying, Fu, Yanwei, Yang, Yujiu, Shao, Jing, Hu, Xia, Liu, Dongrui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
von: Yang, Zhonghao, et al.
Veröffentlicht: (2026)
von: Yang, Zhonghao, et al.
Veröffentlicht: (2026)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
von: Guo, Dadi, et al.
Veröffentlicht: (2025)
von: Guo, Dadi, et al.
Veröffentlicht: (2025)
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
von: Ren, Qihan, et al.
Veröffentlicht: (2026)
von: Ren, Qihan, et al.
Veröffentlicht: (2026)
UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
von: Ding, Yizhuo, et al.
Veröffentlicht: (2025)
von: Ding, Yizhuo, et al.
Veröffentlicht: (2025)
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
von: Guo, Dadi, et al.
Veröffentlicht: (2026)
von: Guo, Dadi, et al.
Veröffentlicht: (2026)
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
von: Shao, Shuai, et al.
Veröffentlicht: (2025)
von: Shao, Shuai, et al.
Veröffentlicht: (2025)
The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
von: Qian, Chen, et al.
Veröffentlicht: (2026)
von: Qian, Chen, et al.
Veröffentlicht: (2026)
Attributing Emergence in Million-Agent Systems
von: Tang, Ling, et al.
Veröffentlicht: (2026)
von: Tang, Ling, et al.
Veröffentlicht: (2026)
COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability
von: Ding, Yizhuo, et al.
Veröffentlicht: (2025)
von: Ding, Yizhuo, et al.
Veröffentlicht: (2025)
Modeling Spatiotemporal Neural Frames for High Resolution Brain Dynamic
von: Qu, Wanying, et al.
Veröffentlicht: (2026)
von: Qu, Wanying, et al.
Veröffentlicht: (2026)
PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning
von: Zhao, Yusong, et al.
Veröffentlicht: (2026)
von: Zhao, Yusong, et al.
Veröffentlicht: (2026)
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
von: Liu, Dongrui, et al.
Veröffentlicht: (2026)
von: Liu, Dongrui, et al.
Veröffentlicht: (2026)
VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
von: Ding, Yizhuo, et al.
Veröffentlicht: (2025)
von: Ding, Yizhuo, et al.
Veröffentlicht: (2025)
Geo-ATBench: A Benchmark for Geospatial Audio Tagging with Geospatial Semantic Context
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
Are Your Agents Upward Deceivers?
von: Guo, Dadi, et al.
Veröffentlicht: (2025)
von: Guo, Dadi, et al.
Veröffentlicht: (2025)
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
von: Liu, Dongrui, et al.
Veröffentlicht: (2026)
von: Liu, Dongrui, et al.
Veröffentlicht: (2026)
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
von: Chen, Guanxu, et al.
Veröffentlicht: (2025)
von: Chen, Guanxu, et al.
Veröffentlicht: (2025)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
von: Hu, Xuhao, et al.
Veröffentlicht: (2024)
von: Hu, Xuhao, et al.
Veröffentlicht: (2024)
Independent characterization of the elastic and the mixing parts of hydrogel osmotic pressure
von: Shao, Zefan, et al.
Veröffentlicht: (2023)
von: Shao, Zefan, et al.
Veröffentlicht: (2023)
LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint
von: Ma, Qianli, et al.
Veröffentlicht: (2025)
von: Ma, Qianli, et al.
Veröffentlicht: (2025)
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
von: Chen, Guanxu, et al.
Veröffentlicht: (2026)
von: Chen, Guanxu, et al.
Veröffentlicht: (2026)
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
von: Chen, Tianyu, et al.
Veröffentlicht: (2026)
von: Chen, Tianyu, et al.
Veröffentlicht: (2026)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
A Trajectory Generator for High-Density Traffic and Diverse Agent-Interaction Scenarios
von: Yang, Ruining, et al.
Veröffentlicht: (2025)
von: Yang, Ruining, et al.
Veröffentlicht: (2025)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
von: Lu, Xiaoya, et al.
Veröffentlicht: (2025)
Interpreting Emergent Extreme Events in Multi-Agent Systems
von: Tang, Ling, et al.
Veröffentlicht: (2026)
von: Tang, Ling, et al.
Veröffentlicht: (2026)
Towards the Dynamics of a DNN Learning Symbolic Interactions
von: Ren, Qihan, et al.
Veröffentlicht: (2024)
von: Ren, Qihan, et al.
Veröffentlicht: (2024)
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
von: Bai, Haoyue, et al.
Veröffentlicht: (2026)
von: Bai, Haoyue, et al.
Veröffentlicht: (2026)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
von: Zhou, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhou, Tianyi, et al.
Veröffentlicht: (2026)
INFA-Guard: Mitigating Malicious Propagation via Infection-Aware Safeguarding in LLM-Based Multi-Agent Systems
von: Zhou, Yijin, et al.
Veröffentlicht: (2026)
von: Zhou, Yijin, et al.
Veröffentlicht: (2026)
LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
von: Chen, Tianyu, et al.
Veröffentlicht: (2026)
von: Chen, Tianyu, et al.
Veröffentlicht: (2026)
MinD-3D: Reconstruct High-quality 3D objects in Human Brain
von: Gao, Jianxiong, et al.
Veröffentlicht: (2023)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2023)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
MinD-3D++: Advancing fMRI-Based 3D Reconstruction with High-Quality Textured Mesh Generation and a Comprehensive Dataset
von: Gao, Jianxiong, et al.
Veröffentlicht: (2024)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2024)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
von: Zhu, Bingwen, et al.
Veröffentlicht: (2026)
Beyond Quantity: Trajectory Diversity Scaling for Code Agents
von: Chen, Guhong, et al.
Veröffentlicht: (2026)
von: Chen, Guhong, et al.
Veröffentlicht: (2026)
PRISM: Preference-Aware Influence Function Based Data Selection Method for Efficient Fine-Tuning
von: Lin, Qihao, et al.
Veröffentlicht: (2026)
von: Lin, Qihao, et al.
Veröffentlicht: (2026)
Tree-frog-inspired osmocapillary adhesive bonding to diverse substrates
von: Shao, Zefan, et al.
Veröffentlicht: (2025)
von: Shao, Zefan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
von: Yang, Zhonghao, et al.
Veröffentlicht: (2026) -
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
von: Yang, Jingyi, et al.
Veröffentlicht: (2025) -
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
von: Guo, Dadi, et al.
Veröffentlicht: (2025) -
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
von: Ren, Qihan, et al.
Veröffentlicht: (2026) -
UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
von: Ding, Yizhuo, et al.
Veröffentlicht: (2025)