No More, No Less: Task Alignment in Terminal Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Mavali, Sina, Pape, David, Evertz, Jonathan, Abedini, Samira, Srivastav, Devansh, Eisenhofer, Thorsten, Abdelnabi, Sahar, Schönherr, Lea |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompt Obfuscation for Large Language Models
by: Pape, David, et al.
Published: (2024)
by: Pape, David, et al.
Published: (2024)
Whispers in the Machine: Confidentiality in Agentic Systems
by: Evertz, Jonathan, et al.
Published: (2024)
by: Evertz, Jonathan, et al.
Published: (2024)
Chasing Shadows: Pitfalls in LLM Security Research
by: Evertz, Jonathan, et al.
Published: (2025)
by: Evertz, Jonathan, et al.
Published: (2025)
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
by: Abdelnabi, Sahar, et al.
Published: (2026)
by: Abdelnabi, Sahar, et al.
Published: (2026)
Don't Trust Stubborn Neighbors: A Security Framework for Agentic Networks
by: Abedini, Samira, et al.
Published: (2026)
by: Abedini, Samira, et al.
Published: (2026)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
by: Salem, Ahmed, et al.
Published: (2026)
by: Salem, Ahmed, et al.
Published: (2026)
Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks
by: De Stefano, Gianluca, et al.
Published: (2024)
by: De Stefano, Gianluca, et al.
Published: (2024)
AI Agents May Always Fall for Prompt Injections
by: Abdelnabi, Sahar, et al.
Published: (2026)
by: Abdelnabi, Sahar, et al.
Published: (2026)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
by: Gomaa, Amr, et al.
Published: (2025)
by: Gomaa, Amr, et al.
Published: (2025)
HexaCoder: Secure Code Generation via Oracle-Guided Synthetic Training Data
by: Hajipour, Hossein, et al.
Published: (2024)
by: Hajipour, Hossein, et al.
Published: (2024)
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
by: Schmotz, David, et al.
Published: (2026)
by: Schmotz, David, et al.
Published: (2026)
Adversarial Observations in Weather Forecasting
by: Imgrund, Erik, et al.
Published: (2025)
by: Imgrund, Erik, et al.
Published: (2025)
Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition
by: Debenedetti, Edoardo, et al.
Published: (2024)
by: Debenedetti, Edoardo, et al.
Published: (2024)
ClawLess: A Security Model of AI Agents
by: Lu, Hongyi, et al.
Published: (2026)
by: Lu, Hongyi, et al.
Published: (2026)
LoopTrap: Termination Poisoning Attacks on LLM Agents
by: Xu, Huiyu, et al.
Published: (2026)
by: Xu, Huiyu, et al.
Published: (2026)
Unknown Unknowns: Why Hidden Intentions in LLMs Evade Detection
by: Srivastav, Devansh, et al.
Published: (2026)
by: Srivastav, Devansh, et al.
Published: (2026)
Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models
by: Zhang, Xiaomei, et al.
Published: (2026)
by: Zhang, Xiaomei, et al.
Published: (2026)
The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility
by: Pape, David, et al.
Published: (2026)
by: Pape, David, et al.
Published: (2026)
Measuring Safety Alignment Effects in Autonomous Security Agents
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
Hardware-Triggered Backdoors
by: Möller, Jonas, et al.
Published: (2026)
by: Möller, Jonas, et al.
Published: (2026)
Less is More: Sparse Watermarking in LLMs with Enhanced Text Quality
by: Hoang, Duy C., et al.
Published: (2024)
by: Hoang, Duy C., et al.
Published: (2024)
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
by: Nakamura, Mason, et al.
Published: (2025)
by: Nakamura, Mason, et al.
Published: (2025)
Get my drift? Catching LLM Task Drift with Activation Deltas
by: Abdelnabi, Sahar, et al.
Published: (2024)
by: Abdelnabi, Sahar, et al.
Published: (2024)
Conning the Crypto Conman: End-to-End Analysis of Cryptocurrency-based Technical Support Scams
by: Acharya, Bhupendra, et al.
Published: (2024)
by: Acharya, Bhupendra, et al.
Published: (2024)
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias
by: Bernstein, Shir, et al.
Published: (2025)
by: Bernstein, Shir, et al.
Published: (2025)
Agent Safety Alignment via Reinforcement Learning
by: Sha, Zeyang, et al.
Published: (2025)
by: Sha, Zeyang, et al.
Published: (2025)
Firewalls to Secure Dynamic LLM Agentic Networks
by: Abdelnabi, Sahar, et al.
Published: (2025)
by: Abdelnabi, Sahar, et al.
Published: (2025)
Less is More: Revisiting the Gaussian Mechanism for Differential Privacy
by: Ji, Tianxi, et al.
Published: (2023)
by: Ji, Tianxi, et al.
Published: (2023)
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
by: Zhao, Lei, et al.
Published: (2026)
by: Zhao, Lei, et al.
Published: (2026)
Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
by: Pulipaka, Sidharth, et al.
Published: (2026)
by: Pulipaka, Sidharth, et al.
Published: (2026)
"That's another doom I haven't thought about": A User Study on AI Labels as a Safeguard Against Image-Based Misinformation
by: Höltervennhoff, Sandra, et al.
Published: (2025)
by: Höltervennhoff, Sandra, et al.
Published: (2025)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Learned-Database Systems Security
by: Schuster, Roei, et al.
Published: (2022)
by: Schuster, Roei, et al.
Published: (2022)
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
by: Fujinuma, Yoshinari, et al.
Published: (2026)
by: Fujinuma, Yoshinari, et al.
Published: (2026)
When Alignment Isn't Enough: Response-Path Attacks on LLM Agents
by: Luo, Mingyu, et al.
Published: (2026)
by: Luo, Mingyu, et al.
Published: (2026)
Agora: Trust Less and Open More in Verification for Confidential Computing
by: Chen, Hongbo, et al.
Published: (2024)
by: Chen, Hongbo, et al.
Published: (2024)
Less Is More: Sparse and Cooperative Perturbation for Point Cloud Attacks
by: Tang, Keke, et al.
Published: (2025)
by: Tang, Keke, et al.
Published: (2025)
Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference
by: Yang, Fan
Published: (2026)
by: Yang, Fan
Published: (2026)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
by: Jia, Feiran, et al.
Published: (2024)
by: Jia, Feiran, et al.
Published: (2024)
Similar Items
-
Prompt Obfuscation for Large Language Models
by: Pape, David, et al.
Published: (2024) -
Whispers in the Machine: Confidentiality in Agentic Systems
by: Evertz, Jonathan, et al.
Published: (2024) -
Chasing Shadows: Pitfalls in LLM Security Research
by: Evertz, Jonathan, et al.
Published: (2025) -
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
by: Abdelnabi, Sahar, et al.
Published: (2026) -
Don't Trust Stubborn Neighbors: A Security Framework for Agentic Networks
by: Abedini, Samira, et al.
Published: (2026)