The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Shasha, Carroll, Fiona, Bentley, Barry L. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment
di: Yu, Shasha, et al.
Pubblicazione: (2026)
di: Yu, Shasha, et al.
Pubblicazione: (2026)
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
di: Castellani, Tommaso, et al.
Pubblicazione: (2025)
di: Castellani, Tommaso, et al.
Pubblicazione: (2025)
Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
di: Zhang, Peng, et al.
Pubblicazione: (2025)
di: Zhang, Peng, et al.
Pubblicazione: (2025)
AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
di: Wang, Zhaohui Geoffrey
Pubblicazione: (2026)
di: Wang, Zhaohui Geoffrey
Pubblicazione: (2026)
SMARLA: A Safety Monitoring Approach for Deep Reinforcement Learning Agents
di: Zolfagharian, Amirhossein, et al.
Pubblicazione: (2023)
di: Zolfagharian, Amirhossein, et al.
Pubblicazione: (2023)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
di: Rank, Ben, et al.
Pubblicazione: (2026)
di: Rank, Ben, et al.
Pubblicazione: (2026)
The Dual-State Architecture for Reliable LLM Agents
di: Thompson, Matthew
Pubblicazione: (2025)
di: Thompson, Matthew
Pubblicazione: (2025)
From Hazard Identification to Controller Design: Proactive and LLM-Supported Safety Engineering for ML-Powered Systems
di: Hong, Yining, et al.
Pubblicazione: (2025)
di: Hong, Yining, et al.
Pubblicazione: (2025)
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
di: Manglik, Akshay, et al.
Pubblicazione: (2026)
di: Manglik, Akshay, et al.
Pubblicazione: (2026)
A Regression Framework for Understanding Prompt Component Impact on LLM Performance
di: Lauziere, Andrew, et al.
Pubblicazione: (2026)
di: Lauziere, Andrew, et al.
Pubblicazione: (2026)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
di: Han, Xiaoke, et al.
Pubblicazione: (2025)
di: Han, Xiaoke, et al.
Pubblicazione: (2025)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
di: Li, Xu, et al.
Pubblicazione: (2026)
di: Li, Xu, et al.
Pubblicazione: (2026)
Causal Fuzzing for Verifying Machine Unlearning
di: Mazhar, Anna, et al.
Pubblicazione: (2025)
di: Mazhar, Anna, et al.
Pubblicazione: (2025)
ToolFactory: Automating Tool Generation by Leveraging LLM to Understand REST API Documentations
di: Ni, Xinyi, et al.
Pubblicazione: (2025)
di: Ni, Xinyi, et al.
Pubblicazione: (2025)
Cleaning Maintenance Logs with LLM Agents for Improved Predictive Maintenance
di: Dimidov, Valeriu, et al.
Pubblicazione: (2025)
di: Dimidov, Valeriu, et al.
Pubblicazione: (2025)
Can Coding Agents Be General Agents?
di: Ivanov, Maksim, et al.
Pubblicazione: (2026)
di: Ivanov, Maksim, et al.
Pubblicazione: (2026)
Standardization Trends on Safety and Trustworthiness Technology for Advanced AI
di: Jeon, Jonghong
Pubblicazione: (2024)
di: Jeon, Jonghong
Pubblicazione: (2024)
Read, Extract, Classify: A Tool for Smarter Requirements Engineering
di: Bhattacharya, Paheli, et al.
Pubblicazione: (2026)
di: Bhattacharya, Paheli, et al.
Pubblicazione: (2026)
MTAD: Tools and Benchmarks for Multivariate Time Series Anomaly Detection
di: Liu, Jinyang, et al.
Pubblicazione: (2024)
di: Liu, Jinyang, et al.
Pubblicazione: (2024)
SMARTCAL: An Approach to Self-Aware Tool-Use Evaluation and Calibration
di: Shen, Yuanhao, et al.
Pubblicazione: (2024)
di: Shen, Yuanhao, et al.
Pubblicazione: (2024)
Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis
di: Zhang, Ke, et al.
Pubblicazione: (2026)
di: Zhang, Ke, et al.
Pubblicazione: (2026)
Instance-Level Safety-Aware Fidelity of Synthetic Data and Its Calibration
di: Cheng, Chih-Hong, et al.
Pubblicazione: (2024)
di: Cheng, Chih-Hong, et al.
Pubblicazione: (2024)
AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks
di: Oulkadda, Ilyas, et al.
Pubblicazione: (2025)
di: Oulkadda, Ilyas, et al.
Pubblicazione: (2025)
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
di: Feng, Yunfei, et al.
Pubblicazione: (2026)
di: Feng, Yunfei, et al.
Pubblicazione: (2026)
VibeTensor: System Software for Deep Learning, Fully Generated by AI Agents
di: Xu, Bing, et al.
Pubblicazione: (2026)
di: Xu, Bing, et al.
Pubblicazione: (2026)
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
di: Bisztray, Tamas, et al.
Pubblicazione: (2025)
di: Bisztray, Tamas, et al.
Pubblicazione: (2025)
Understanding LLM-Driven Test Oracle Generation
di: Bodicoat, Adam, et al.
Pubblicazione: (2026)
di: Bodicoat, Adam, et al.
Pubblicazione: (2026)
A Survey on Code Generation with LLM-based Agents
di: Dong, Yihong, et al.
Pubblicazione: (2025)
di: Dong, Yihong, et al.
Pubblicazione: (2025)
Agentless: Demystifying LLM-based Software Engineering Agents
di: Xia, Chunqiu Steven, et al.
Pubblicazione: (2024)
di: Xia, Chunqiu Steven, et al.
Pubblicazione: (2024)
The BrowserGym Ecosystem for Web Agent Research
di: De Chezelles, Thibault Le Sellier, et al.
Pubblicazione: (2024)
di: De Chezelles, Thibault Le Sellier, et al.
Pubblicazione: (2024)
Mutation-Guided LLM-based Test Generation at Meta
di: Foster, Christopher, et al.
Pubblicazione: (2025)
di: Foster, Christopher, et al.
Pubblicazione: (2025)
Retrieval-Augmented Instruction Tuning for Automated Process Engineering Calculations : A Tool-Chaining Problem-Solving Framework with Attributable Reflection
di: Sakhinana, Sagar Srinivas, et al.
Pubblicazione: (2024)
di: Sakhinana, Sagar Srinivas, et al.
Pubblicazione: (2024)
SWE-Bench-CL: Continual Learning for Coding Agents
di: Joshi, Thomas, et al.
Pubblicazione: (2025)
di: Joshi, Thomas, et al.
Pubblicazione: (2025)
Automated Cloud Infrastructure-as-Code Reconciliation with AI Agents
di: Yang, Zhenning, et al.
Pubblicazione: (2025)
di: Yang, Zhenning, et al.
Pubblicazione: (2025)
LLMs in Coding and their Impact on the Commercial Software Engineering Landscape
di: Belozerov, Vladislav, et al.
Pubblicazione: (2025)
di: Belozerov, Vladislav, et al.
Pubblicazione: (2025)
The Impact of Environment Configurations on the Stability of AI-Enabled Systems
di: Rahman, Musfiqur, et al.
Pubblicazione: (2024)
di: Rahman, Musfiqur, et al.
Pubblicazione: (2024)
On The Impact of Merge Request Deviations on Code Review Practices
di: Kansab, Samah, et al.
Pubblicazione: (2025)
di: Kansab, Samah, et al.
Pubblicazione: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
di: Xu, WeiZhe, et al.
Pubblicazione: (2026)
di: Xu, WeiZhe, et al.
Pubblicazione: (2026)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
di: Jiang, Shan, et al.
Pubblicazione: (2026)
di: Jiang, Shan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment
di: Yu, Shasha, et al.
Pubblicazione: (2026) -
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
di: Castellani, Tommaso, et al.
Pubblicazione: (2025) -
Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
di: Zhang, Peng, et al.
Pubblicazione: (2025) -
AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems
di: Wang, Zhaohui Geoffrey
Pubblicazione: (2026) -
SMARLA: A Safety Monitoring Approach for Deep Reinforcement Learning Agents
di: Zolfagharian, Amirhossein, et al.
Pubblicazione: (2023)