TFL: Targeted Bit-Flip Attack on Large Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Jingkai, Chakrabarti, Chaitali, Fan, Deliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
von: Guo, Jingkai, et al.
Veröffentlicht: (2025)
von: Guo, Jingkai, et al.
Veröffentlicht: (2025)
PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips
von: Coalson, Zachary, et al.
Veröffentlicht: (2024)
von: Coalson, Zachary, et al.
Veröffentlicht: (2024)
SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
von: Xu, Haotian, et al.
Veröffentlicht: (2025)
von: Xu, Haotian, et al.
Veröffentlicht: (2025)
User Inference Attacks on Large Language Models
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
von: He, Jiaming, et al.
Veröffentlicht: (2024)
von: He, Jiaming, et al.
Veröffentlicht: (2024)
An Early Categorization of Prompt Injection Attacks on Large Language Models
von: Rossi, Sippo, et al.
Veröffentlicht: (2024)
von: Rossi, Sippo, et al.
Veröffentlicht: (2024)
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
Confidence Elicitation: A New Attack Vector for Large Language Models
von: Formento, Brian, et al.
Veröffentlicht: (2025)
von: Formento, Brian, et al.
Veröffentlicht: (2025)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025)
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
von: Sahili, Ali Al, et al.
Veröffentlicht: (2025)
von: Sahili, Ali Al, et al.
Veröffentlicht: (2025)
Cyber-Attack Technique Classification Using Two-Stage Trained Large Language Models
von: You, Weiqiu, et al.
Veröffentlicht: (2024)
von: You, Weiqiu, et al.
Veröffentlicht: (2024)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models
von: Wu, Yihan, et al.
Veröffentlicht: (2023)
von: Wu, Yihan, et al.
Veröffentlicht: (2023)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
von: Xue, Eric, et al.
Veröffentlicht: (2025)
von: Xue, Eric, et al.
Veröffentlicht: (2025)
Cross-Entropy Attacks to Language Models via Rare Event Simulation
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
Impactful Bit-Flip Search on Full-precision Models
von: Benedek, Nadav, et al.
Veröffentlicht: (2024)
von: Benedek, Nadav, et al.
Veröffentlicht: (2024)
Model Provenance Testing for Large Language Models
von: Nikolic, Ivica, et al.
Veröffentlicht: (2025)
von: Nikolic, Ivica, et al.
Veröffentlicht: (2025)
A Watermark for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
On the Reliability of Watermarks for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Verification of Bit-Flip Attacks against Quantized Neural Networks
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
A Transfer Attack to Image Watermarks
von: Hu, Yuepeng, et al.
Veröffentlicht: (2024)
von: Hu, Yuepeng, et al.
Veröffentlicht: (2024)
Topic-Based Watermarks for Large Language Models
von: Nemecek, Alexander, et al.
Veröffentlicht: (2024)
von: Nemecek, Alexander, et al.
Veröffentlicht: (2024)
Representation Bending for Large Language Model Safety
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
SPML: A DSL for Defending Language Models Against Prompt Attacks
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2024)
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2024)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
Detecting Pretraining Data from Large Language Models
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
Learning to Poison Large Language Models for Downstream Manipulation
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2024)
Watermarks for Embeddings-as-a-Service Large Language Models
von: Shetty, Anudeex
Veröffentlicht: (2025)
von: Shetty, Anudeex
Veröffentlicht: (2025)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
Membership Inference Attacks and Privacy in Topic Modeling
von: Manzonelli, Nico, et al.
Veröffentlicht: (2024)
von: Manzonelli, Nico, et al.
Veröffentlicht: (2024)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
von: Roth, Tom, et al.
Veröffentlicht: (2021)
von: Roth, Tom, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
von: Guo, Jingkai, et al.
Veröffentlicht: (2025) -
PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips
von: Coalson, Zachary, et al.
Veröffentlicht: (2024) -
SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
von: Xu, Haotian, et al.
Veröffentlicht: (2025) -
User Inference Attacks on Large Language Models
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023) -
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)