Claim-Guided Textual Backdoor Attack for Practical Applications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Minkyoo, Kim, Hanna, Kim, Jaehan, Jin, Youngjin, Shin, Seungwon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Obliviate: Neutralizing Task-agnostic Backdoors within the Parameter-efficient Fine-tuning Paradigm
von: Kim, Jaehan, et al.
Veröffentlicht: (2024)
von: Kim, Jaehan, et al.
Veröffentlicht: (2024)
PassREfinder-FL: Privacy-Preserving Credential Stuffing Risk Prediction via Graph-Based Federated Learning for Representing Password Reuse between Websites
von: Kim, Jaehan, et al.
Veröffentlicht: (2025)
von: Kim, Jaehan, et al.
Veröffentlicht: (2025)
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
von: Kim, Jaehan, et al.
Veröffentlicht: (2025)
von: Kim, Jaehan, et al.
Veröffentlicht: (2025)
Subgraph Reconstruction Attacks on Graph RAG Deployments with Practical Defenses
von: Song, Minkyoo, et al.
Veröffentlicht: (2026)
von: Song, Minkyoo, et al.
Veröffentlicht: (2026)
When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs
von: Kim, Hanna, et al.
Veröffentlicht: (2024)
von: Kim, Hanna, et al.
Veröffentlicht: (2024)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
von: Du, Wei, et al.
Veröffentlicht: (2023)
von: Du, Wei, et al.
Veröffentlicht: (2023)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Invisible Textual Backdoor Attacks based on Dual-Trigger
von: Hou, Yang, et al.
Veröffentlicht: (2024)
von: Hou, Yang, et al.
Veröffentlicht: (2024)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
RedacBench: Can AI Erase Your Secrets?
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
Activation-Guided Local Editing for Jailbreaking Attacks
von: Wang, Jiecong, et al.
Veröffentlicht: (2025)
von: Wang, Jiecong, et al.
Veröffentlicht: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
An Investigation on Group Query Hallucination Attacks
von: Miao, Kehao, et al.
Veröffentlicht: (2025)
von: Miao, Kehao, et al.
Veröffentlicht: (2025)
Exploring Backdoor Vulnerabilities of Chat Models
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2024)
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Goal-guided Generative Prompt Injection Attack on Large Language Models
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
von: Ji, Wence, et al.
Veröffentlicht: (2025)
von: Ji, Wence, et al.
Veröffentlicht: (2025)
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
von: Hu, Man, et al.
Veröffentlicht: (2025)
von: Hu, Man, et al.
Veröffentlicht: (2025)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
von: Cao, Yuanpu, et al.
Veröffentlicht: (2023)
von: Cao, Yuanpu, et al.
Veröffentlicht: (2023)
LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem
von: Liu, Hongyi, et al.
Veröffentlicht: (2024)
von: Liu, Hongyi, et al.
Veröffentlicht: (2024)
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
von: Wang, Xilong, et al.
Veröffentlicht: (2026)
von: Wang, Xilong, et al.
Veröffentlicht: (2026)
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
von: Lu, Ning, et al.
Veröffentlicht: (2023)
von: Lu, Ning, et al.
Veröffentlicht: (2023)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
von: Yi, Sibo, et al.
Veröffentlicht: (2025)
von: Yi, Sibo, et al.
Veröffentlicht: (2025)
Concept-Guided Backdoor Attack on Vision Language Models
von: Shen, Haoyu, et al.
Veröffentlicht: (2025)
von: Shen, Haoyu, et al.
Veröffentlicht: (2025)
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
von: Yu, Ye, et al.
Veröffentlicht: (2026)
von: Yu, Ye, et al.
Veröffentlicht: (2026)
HauntAttack: When Attack Follows Reasoning as a Shadow
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
von: Ma, Jingyuan, et al.
Veröffentlicht: (2025)
Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game
von: Xie, Yuanbo, et al.
Veröffentlicht: (2026)
von: Xie, Yuanbo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Obliviate: Neutralizing Task-agnostic Backdoors within the Parameter-efficient Fine-tuning Paradigm
von: Kim, Jaehan, et al.
Veröffentlicht: (2024) -
PassREfinder-FL: Privacy-Preserving Credential Stuffing Risk Prediction via Graph-Based Federated Learning for Representing Password Reuse between Websites
von: Kim, Jaehan, et al.
Veröffentlicht: (2025) -
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
von: Kim, Jaehan, et al.
Veröffentlicht: (2025) -
Subgraph Reconstruction Attacks on Graph RAG Deployments with Practical Defenses
von: Song, Minkyoo, et al.
Veröffentlicht: (2026) -
When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs
von: Kim, Hanna, et al.
Veröffentlicht: (2024)