A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mo, Lingbo, Liao, Zeyi, Zheng, Boyuan, Su, Yu, Xiao, Chaowei, Sun, Huan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
von: Jones, Jaylen, et al.
Veröffentlicht: (2024)
von: Jones, Jaylen, et al.
Veröffentlicht: (2024)
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
von: Mo, Lingbo, et al.
Veröffentlicht: (2023)
von: Mo, Lingbo, et al.
Veröffentlicht: (2023)
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
von: Zhang, Kai, et al.
Veröffentlicht: (2023)
von: Zhang, Kai, et al.
Veröffentlicht: (2023)
WebGuard: Building a Generalizable Guardrail for Web Agents
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
von: Luo, Weidi, et al.
Veröffentlicht: (2024)
von: Luo, Weidi, et al.
Veröffentlicht: (2024)
Fast Adversarial Training against Textual Adversarial Attacks
von: Yang, Yichen, et al.
Veröffentlicht: (2024)
von: Yang, Yichen, et al.
Veröffentlicht: (2024)
AttributionBench: How Hard is Automatic Attribution Evaluation?
von: Li, Yifei, et al.
Veröffentlicht: (2024)
von: Li, Yifei, et al.
Veröffentlicht: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
von: Li, Hao, et al.
Veröffentlicht: (2026)
von: Li, Hao, et al.
Veröffentlicht: (2026)
House of Cards: Massive Weights in LLMs
von: Oh, Jaehoon, et al.
Veröffentlicht: (2024)
von: Oh, Jaehoon, et al.
Veröffentlicht: (2024)
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
von: Gou, Boyu, et al.
Veröffentlicht: (2024)
von: Gou, Boyu, et al.
Veröffentlicht: (2024)
Combating Adversarial Attacks with Multi-Agent Debate
von: Chern, Steffi, et al.
Veröffentlicht: (2024)
von: Chern, Steffi, et al.
Veröffentlicht: (2024)
AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
von: Shu, Yiheng, et al.
Veröffentlicht: (2026)
von: Shu, Yiheng, et al.
Veröffentlicht: (2026)
AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
von: Gou, Boyu, et al.
Veröffentlicht: (2025)
von: Gou, Boyu, et al.
Veröffentlicht: (2025)
HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals
von: Mo, Lingbo, et al.
Veröffentlicht: (2024)
von: Mo, Lingbo, et al.
Veröffentlicht: (2024)
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2023)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2023)
ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
von: Chen, Ziru, et al.
Veröffentlicht: (2024)
von: Chen, Ziru, et al.
Veröffentlicht: (2024)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
von: Xu, Chejian, et al.
Veröffentlicht: (2024)
MultiAgent Collaboration Attack: Investigating Adversarial Attacks in Large Language Model Collaborations via Debate
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2024)
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2024)
Exploring Gradient-Guided Masked Language Model to Detect Textual Adversarial Attacks
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
Semantic Representation Attack against Aligned Large Language Models
von: Lian, Jiawei, et al.
Veröffentlicht: (2025)
von: Lian, Jiawei, et al.
Veröffentlicht: (2025)
Preference Poisoning Attacks on Reward Model Learning
von: Wu, Junlin, et al.
Veröffentlicht: (2024)
von: Wu, Junlin, et al.
Veröffentlicht: (2024)
DROJ: A Prompt-Driven Attack against Large Language Models
von: Hu, Leyang, et al.
Veröffentlicht: (2024)
von: Hu, Leyang, et al.
Veröffentlicht: (2024)
SEP-Attack: A Simple and Effective Paradigm for Transfer-Based Textual Adversarial Attack
von: Liu, Han, et al.
Veröffentlicht: (2026)
von: Liu, Han, et al.
Veröffentlicht: (2026)
Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
von: Chen, Yurun, et al.
Veröffentlicht: (2025)
von: Chen, Yurun, et al.
Veröffentlicht: (2025)
Round Trip Translation Defence against Large Language Model Jailbreaking Attacks
von: Yung, Canaan, et al.
Veröffentlicht: (2024)
von: Yung, Canaan, et al.
Veröffentlicht: (2024)
HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on Text
von: Liu, Han, et al.
Veröffentlicht: (2024)
von: Liu, Han, et al.
Veröffentlicht: (2024)
Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks
von: Lin, Tzu-Ling, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Ling, et al.
Veröffentlicht: (2025)
An Illusion of Progress? Assessing the Current State of Web Agents
von: Xue, Tianci, et al.
Veröffentlicht: (2025)
von: Xue, Tianci, et al.
Veröffentlicht: (2025)
SafePred: A Predictive Guardrail for Computer-Using Agents via World Models
von: Chen, Yurun, et al.
Veröffentlicht: (2026)
von: Chen, Yurun, et al.
Veröffentlicht: (2026)
SoLA: Leveraging Soft Activation Sparsity and Low-Rank Decomposition for Large Language Model Compression
von: Huang, Xinhao, et al.
Veröffentlicht: (2026)
von: Huang, Xinhao, et al.
Veröffentlicht: (2026)
$\textit{LinkPrompt}$: Natural and Universal Adversarial Attacks on Prompt-based Language Models
von: Xu, Yue, et al.
Veröffentlicht: (2024)
von: Xu, Yue, et al.
Veröffentlicht: (2024)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
von: Zhou, Weibo, et al.
Veröffentlicht: (2025)
von: Zhou, Weibo, et al.
Veröffentlicht: (2025)
What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction
von: Li, Lingbo, et al.
Veröffentlicht: (2025)
von: Li, Lingbo, et al.
Veröffentlicht: (2025)
Revisiting Character-level Adversarial Attacks for Language Models
von: Rocamora, Elias Abad, et al.
Veröffentlicht: (2024)
von: Rocamora, Elias Abad, et al.
Veröffentlicht: (2024)
Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
SudoLM: Learning Access Control of Parametric Knowledge with Authorization Alignment
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
von: Liao, Zeyi, et al.
Veröffentlicht: (2024) -
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
von: Jones, Jaylen, et al.
Veröffentlicht: (2024) -
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
von: Mo, Lingbo, et al.
Veröffentlicht: (2023) -
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
von: Zhang, Kai, et al.
Veröffentlicht: (2023) -
WebGuard: Building a Generalizable Guardrail for Web Agents
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)