Tool Preferences in Agentic LLMs are Unreliable
Fuente:
arXiv
Salvato in:
| Autori principali: | Faghih, Kazem, Wang, Wenxiao, Cheng, Yize, Bharti, Siddhant, Sriramanan, Gaurang, Balasubramanian, Sriram, Hosseini, Parsa, Feizi, Soheil |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fast Adversarial Attacks on Language Models In One GPU Minute
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2024)
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2024)
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
di: Saha, Shoumik, et al.
Pubblicazione: (2026)
di: Saha, Shoumik, et al.
Pubblicazione: (2026)
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
di: Wang, Wenxiao, et al.
Pubblicazione: (2025)
di: Wang, Wenxiao, et al.
Pubblicazione: (2025)
Certifying LLM Safety against Adversarial Prompting
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception
di: Cheng, Yize, et al.
Pubblicazione: (2025)
di: Cheng, Yize, et al.
Pubblicazione: (2025)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
di: Wang, Che, et al.
Pubblicazione: (2026)
di: Wang, Che, et al.
Pubblicazione: (2026)
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
di: Panpatil, Siddhant, et al.
Pubblicazione: (2025)
di: Panpatil, Siddhant, et al.
Pubblicazione: (2025)
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
di: Mohseni, Seyedreza, et al.
Pubblicazione: (2024)
di: Mohseni, Seyedreza, et al.
Pubblicazione: (2024)
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2025)
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2025)
Quantitative Certification of Agentic Tool Selection
di: Yeon, Jehyeok, et al.
Pubblicazione: (2025)
di: Yeon, Jehyeok, et al.
Pubblicazione: (2025)
Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools
di: Mohammadi, Bardia, et al.
Pubblicazione: (2026)
di: Mohammadi, Bardia, et al.
Pubblicazione: (2026)
Protecting Your LLMs with Information Bottleneck
di: Liu, Zichuan, et al.
Pubblicazione: (2024)
di: Liu, Zichuan, et al.
Pubblicazione: (2024)
Can AI-Generated Text be Reliably Detected?
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2023)
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2023)
A Survey on Agentic Security: Applications, Threats and Defenses
di: Shahriar, Asif, et al.
Pubblicazione: (2025)
di: Shahriar, Asif, et al.
Pubblicazione: (2025)
SastBench: A Benchmark for Testing Agentic SAST Triage
di: Feiglin, Jake, et al.
Pubblicazione: (2026)
di: Feiglin, Jake, et al.
Pubblicazione: (2026)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
di: Shen, Guangyu, et al.
Pubblicazione: (2024)
di: Shen, Guangyu, et al.
Pubblicazione: (2024)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
di: Ding, Ruomeng, et al.
Pubblicazione: (2026)
di: Ding, Ruomeng, et al.
Pubblicazione: (2026)
LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
di: Li, Xingyu, et al.
Pubblicazione: (2025)
di: Li, Xingyu, et al.
Pubblicazione: (2025)
AgenTRIM: Tool Risk Mitigation for Agentic AI
di: Betser, Roy, et al.
Pubblicazione: (2026)
di: Betser, Roy, et al.
Pubblicazione: (2026)
AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation
di: Roy, Joyjit, et al.
Pubblicazione: (2026)
di: Roy, Joyjit, et al.
Pubblicazione: (2026)
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models
di: Belkhiter, Yannis, et al.
Pubblicazione: (2026)
di: Belkhiter, Yannis, et al.
Pubblicazione: (2026)
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
di: Kim, Jinhwa, et al.
Pubblicazione: (2025)
di: Kim, Jinhwa, et al.
Pubblicazione: (2025)
AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models
di: Zhang, Jinchuan, et al.
Pubblicazione: (2025)
di: Zhang, Jinchuan, et al.
Pubblicazione: (2025)
Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2025)
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2025)
Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection
di: Chen, Jiaqi, et al.
Pubblicazione: (2024)
di: Chen, Jiaqi, et al.
Pubblicazione: (2024)
Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
di: Qiao, Yuxuan, et al.
Pubblicazione: (2025)
di: Qiao, Yuxuan, et al.
Pubblicazione: (2025)
Defend LLMs Through Self-Consciousness
di: Huang, Boshi, et al.
Pubblicazione: (2025)
di: Huang, Boshi, et al.
Pubblicazione: (2025)
AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
di: Chen, Jizhou, et al.
Pubblicazione: (2025)
di: Chen, Jizhou, et al.
Pubblicazione: (2025)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
di: Chae, Kyubyung, et al.
Pubblicazione: (2025)
di: Chae, Kyubyung, et al.
Pubblicazione: (2025)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
di: Xu, Zhao, et al.
Pubblicazione: (2024)
di: Xu, Zhao, et al.
Pubblicazione: (2024)
The Ethics of Interaction: Mitigating Security Threats in LLMs
di: Kumar, Ashutosh, et al.
Pubblicazione: (2024)
di: Kumar, Ashutosh, et al.
Pubblicazione: (2024)
Strategic Deflection: Defending LLMs from Logit Manipulation
di: Rachidy, Yassine, et al.
Pubblicazione: (2025)
di: Rachidy, Yassine, et al.
Pubblicazione: (2025)
PARASITE: Conditional System Prompt Poisoning to Hijack LLMs
di: Pham, Viet, et al.
Pubblicazione: (2025)
di: Pham, Viet, et al.
Pubblicazione: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
di: Liu, Fan, et al.
Pubblicazione: (2024)
di: Liu, Fan, et al.
Pubblicazione: (2024)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
di: Akbar-Tajari, Mohammad, et al.
Pubblicazione: (2025)
di: Akbar-Tajari, Mohammad, et al.
Pubblicazione: (2025)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
di: Pu, Rui, et al.
Pubblicazione: (2024)
di: Pu, Rui, et al.
Pubblicazione: (2024)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
di: Yang, Yan, et al.
Pubblicazione: (2024)
di: Yang, Yan, et al.
Pubblicazione: (2024)
Waterfall: Framework for Robust and Scalable Text Watermarking and Provenance for LLMs
di: Lau, Gregory Kang Ruey, et al.
Pubblicazione: (2024)
di: Lau, Gregory Kang Ruey, et al.
Pubblicazione: (2024)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
di: Ji, Wence, et al.
Pubblicazione: (2025)
di: Ji, Wence, et al.
Pubblicazione: (2025)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
di: Liu, Songyang, et al.
Pubblicazione: (2025)
di: Liu, Songyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Fast Adversarial Attacks on Language Models In One GPU Minute
di: Sadasivan, Vinu Sankar, et al.
Pubblicazione: (2024) -
Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
di: Saha, Shoumik, et al.
Pubblicazione: (2026) -
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
di: Wang, Wenxiao, et al.
Pubblicazione: (2025) -
Certifying LLM Safety against Adversarial Prompting
di: Kumar, Aounon, et al.
Pubblicazione: (2023) -
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception
di: Cheng, Yize, et al.
Pubblicazione: (2025)